DataEyesAI
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
  1. Wan Video Generation
  • Getting Started
    • Overview
    • Console (Getting Started)
    • API Key
    • Base URL
  • Developer Tool Integration
    • OpenClaw
    • Claude Code
    • Codex
    • Gemini CLI
    • Grok CLI
    • Other Tools
  • AI Models API
    • OpenAI format (supports major original models)
      • Chat (Response)
        • Create Network Search
        • Create Model Response GPT-5 Enable Thinking
        • Create Function Call
        • Create Model Response
        • Create Model Response (Streaming Return)
        • Create Model Response (Control Thinking Length)
      • ChatGPT Interface
        • Audio
          • Audio to text gpt-4o-transcribe
          • GPT-4o-audio
          • Audio to text whisper-1
          • Audio to text gpt-4o-transcribe
          • Create voice gpt-4o-mini-tts
        • Chat
          • Create chat-based image recognition (non-streaming)
          • Create chat-based image recognition (streaming)
          • Create chat-based image recognition (streaming) best64
          • Official N test
          • Create structured output
          • Control the effort level of the inference model
          • Create chat function call
          • deepseek-ocr recognition
          • Create chat completion (non-stream)
        • Completions
          • ChatGPT automatic completion
          • Create completion
      • Image
        • Edit image
        • Create chat completion (streaming)
        • Create chat completion (qwen-mt-turbo)
        • Create chat completion with deepseek v3.1 level of reasoning (streaming)
      • Audio
        • Speech recognition
        • Speech synthesis
        • Official Function Calling invocation
        • Create chat-generated images (non-streaming)
      • Embedding
        • Text embeddings
    • Anthropic format
      • Chat
      • Chat(prompt cache)
      • Streaming response
      • Chat (deep reasoning)
      • Tool invocation (function call)
      • Analyze image
    • Google Gemini interface
      • Native format
        • Text-to-image + control over aspect ratio + clarity
        • Generate image
        • Text generation
        • Text generation - stream
        • Text generation + reasoning - stream
        • Image generation
        • Formatted output
        • Function call
        • Document understanding
        • URL context [native format]
        • Code execution
        • Video understanding
        • URL context
        • Video understanding - url [native format]
        • Imagen 4
        • Audio understanding
        • Embeddings
        • Chat
        • Edit image
      • Image-to-image Base64 request method
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
        • Image editing
        • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image generation( gemini-2.5-flash-image)
        • Image generation gemini-2.5-flash-image, controlling aspect ratio.
        • Image understanding
      • Image-to-image URL request returns URL request format OpenAI
        • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image understanding
    • NanoBanana
      • OpenAI request
        • Edit image
        • OpenAI image format
      • Gemini request
        • Generate image
        • Edit image
    • Midjourney format
      • Midjourney API Reference
      • Task query interface
      • Upload image
      • Get seed (Seed)
      • Submit Imagine task
      • Query tasks based on ID list
      • FaceSwap
      • Execute Action operation
      • /mj/submit/blend
      • Submit Describe task
      • Submit Modal
      • Refresh link
      • Edit image
      • Query task status by task ID
      • Get the seed of the task image
    • Doubao - Painting
      • doubao-seededit-3-0-i2i-250628
      • doubao-seedream-4-0-250828 - text-to-image
      • doubao-seedream-4-0-250828 - image-to-image
      • doubao-seedream-4-0-250828 - multi-image generation
    • Rerank Reordering Model
      • Rerank
    • Video Model
      • Grok Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-Video-Extension
      • Seedance Video Generation
        • 00-Overview
        • 01-Create-Video-Generation-Task
        • 02-Query-Video-Generation-Task
        • 03-Query-Video-Generation-Task-List
        • 04-Cancel-or-Delete-Task
        • Seedance Private Asset Library API Documentation
      • MiniMax-H3 Video Generation
        • 00-Overview
        • 01-Create-Video-Generation
        • 02-Create-Video-Regeneration
        • 03-Create-H3-Context-IR
        • 04-Query-Task
        • 05-List-Tasks
        • 06-Cancel-or-Delete-Task
      • Hailuo Video Generation
        • 00-Overview
        • 01-Text-to-Video-T2V
        • 02-Image-to-Video-I2V
        • 03-First-Last-Frame-FL2V
        • 04-Subject-Reference-S2V
        • 05-Query-Task-Status
        • 06-Video-Download
        • 99-Appendix-Camera-Movement-and-Webhooks
      • Jimeng Video Generation
        • 00-Overview
        • 01-3.0-Pro-Video-Generation
        • 02-720P-Text-to-Video
        • 03-720P-Image-to-Video-First-Frame
        • 04-720P-Image-to-Video-Start-End-Frame
        • 05-720P-Image-to-Video-Camera
        • 06-1080P-Text-to-Video
        • 07-1080P-Image-to-Video-First-Frame
        • 08-1080P-Image-to-Video-Start-End-Frame
        • 09-Error-Codes
      • Kling AI Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Omni-Video
        • 04-Multi-Image-to-Video
        • 05-Motion-Control
        • 06-Multi-Elements
        • 07-Video-Extension
        • 08-Lip-Sync
        • 09-Avatar
        • 10-Text-to-Audio
        • 11-Video-to-Audio
        • 12-TTS
        • 13-Custom-Voices
        • 14-Image-Recognition
        • 15-Element-Management
        • 16-Video-Effects
      • Vidu Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Start-End-Frame
        • 05-Multi-Frame
        • 06-Scene-Template
        • 07-Template-Story
        • 08-Query-Tasks
      • HappyHorse
        • HappyHorse Text-to-Video
        • HappyHorse Image-to-Video (First Frame)
        • HappyHorse Reference-to-Video
        • HappyHorse Video Editing
      • Wan Video Generation
        • 00-Overview.md
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-First-Last-Frame-to-Video
        • 06-Motion-Transfer-and-Character-Swap
        • 07-Digital-Human-Video
        • 08-VACE-Video-Editing
        • 09-Query-Task
    • Audio API
      • Audio API
      • Gemini TTS API
      • Google DeepMind Lyria API
      • Elevenlabs Speech to Text API Reference
      • Text-to-Music Suno
        • Task Submission
          • Generate Song (Inspiration Mode)
          • Generate Song (Custom Mode)
          • Generate Song (Continuation Mode)
          • Generate Song (Singer Style)
          • Generate Song (Secondary Creation from Uploaded Song)
          • Generate Song (Song Stitching)
          • Generate Lyrics
          • Song Stitching
        • Query Interface
          • Batch Retrieve Tasks
          • Query Single Task
  • Search / Reader Product
    • Web Reader API​​
      • Web Reader API
      • Web Reader API(HK)
    • Web Search API​​
      • Modal Card API
        • Weather
          • All City ID
          • Weather Query API
      • Web Search API
      • Video Search api
      • Trending Search API
    • Document OCR Parsing API
      • fiel upload
      • URL Parsing
  • Advanced & System API
    • Data Updates
    • System interface
      • API Key & Quota Query API
      • API Key Management API
    • API Reference​​
      • Error Codes
      • HTTP Notes
    • List models
      • Models
  1. Wan Video Generation

08-VACE-Video-Editing

VACE Unified Video Editing#

Doc version:v1.0.0 | Last updated:2026-07-22
This platform fully supports the official Tongyi Wanxiang (Wan) video generation API. Requests and responses are transparently proxied; parameter semantics are identical to the official API.
The Wan 2.1 unified video editing model (wanx2.1-vace-plus) supports multimodal input — text / image / video — and switches between five editing capabilities within a single endpoint via input.function. Tasks are asynchronous: creating a task returns output.task_id, then poll the query endpoint for results. Generation typically takes about 5–10 minutes.

Capability Overview#

functionCapabilityKey inputs
image_referenceMulti-Image Reference to videoref_images_url (1–3 reference images, subject/background) + prompt; blends them into a coherent video
video_repaintingVideo Repaintingvideo_url; extracts subject motion / composition contours / line-art structure (control_condition) and regenerates; optionally add 1 reference image to replace the subject
video_editVideo Edit (inpainting)video_url + mask (mask_image_url / mask_video_url, one of the two); add / modify / remove elements in the specified region
video_extensionVideo ExtensionFirst/last frame images (first_frame_url / last_frame_url) or first/last video clips (first_clip_url / last_clip_url); generates continuation content
video_outpaintingVideo Outpaintingvideo_url + expansion ratios in four directions (top_scale, etc.); extends content beyond the frame

Output Specifications#

ItemRule
DurationFixed at 5 seconds (duration is fixed to 5 and cannot be changed). For Video Repainting / Video Edit / Video Outpainting: output duration = input duration, up to 5 seconds (3s input → 3s output; 6s input → first 5s output)
Resolution720P tier only. For image_reference, specified via size (default 1280*720); for video-input capabilities: inputs ≤720P keep their original resolution, inputs >720P are scaled down by aspect ratio to at most 720P
FormatMP4 (H.264 encoding); the download URL is valid for 24 hours — save the file promptly
AudioOnly silent video is generated; audio output is not supported
video_extension duration note (important): The output of Video Extension has a total duration of 5 seconds, meaning "input clip + newly generated content" together total 5 seconds — it is not 5 extra seconds appended to the original video. For example, with a 3-second first clip, only about 2 seconds of continuation content is newly generated.

Create Task#

POST https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis

Request Headers#

ParameterTypeRequiredDefaultDescription
Content-TypestringYesapplication/jsonData exchange format
AuthorizationstringYesAuthentication, Bearer {API_KEY}
The X-DashScope-Async: enable header required by the official API is added automatically by the platform — you do not need to include it.

Common Parameters#

The following parameters are shared by all functions:
ParameterTypeRequiredDefaultDescription
modelstringYesModel identifier, fixed to wanx2.1-vace-plus
inputobjectYesBasic input information
input.functionstringYesCapability switch: image_reference, video_repainting, video_edit, video_extension, video_outpainting
input.promptstringYesText prompt, in Chinese or English, length ≤ 800 characters; longer prompts are automatically truncated
parametersobjectDepends on functionProcessing parameters; required for video_repainting (must include control_condition), optional for the others
parameters.durationintegerNo5Output video duration in seconds, fixed at 5, cannot be modified
parameters.prompt_extendbooleanNotrueWhether to enable intelligent prompt rewriting. Recommended to set to false for capabilities with video input (rewriting can misinterpret when the text and video content diverge)
parameters.seedintegerNoRandomRandom seed, range [0, 2147483647]; the same seed yields relatively stable results
parameters.watermarkbooleanNofalseWhether to add an "AI generated" watermark in the bottom-right corner of the video
Function-specific fields are listed below per function.

image_reference (Multi-Image Reference to video)#

input-specific fields:
ParameterTypeRequiredDefaultDescription
input.ref_images_urlarray[string]YesArray of reference image URLs, 1–3 images (only the first 3 are used if more are provided). Recommendations: each subject image should contain only one subject on a solid-color background; at most 1 background image, containing no subject
parameters-specific fields:
ParameterTypeRequiredDefaultDescription
parameters.obj_or_bgarray[string]No["obj"] for a single imageIdentifies the purpose of each image, in one-to-one correspondence with ref_images_url: obj = subject image / bg = background image (at most 1 bg). Explicitly passing it is recommended; its length must match ref_images_url, otherwise an error is returned. Example: ["obj","obj","bg"]
parameters.sizestringNo1280*720Output resolution (720P tier only): 1280*720(16:9), 720*1280(9:16), 960*960(1:1), 832*1088(3:4), 1088*832(4:3)

video_repainting (Video Repainting)#

input-specific fields:
ParameterTypeRequiredDefaultDescription
input.video_urlstringYesInput video URL; limits are described in "Common Input Material Limits"
input.ref_images_urlarray[string]No1 image only, preferably a subject image, used to replace the subject in the video (swap the character while keeping the motion)
parameters-specific fields (parameters is required for this capability):
ParameterTypeRequiredDefaultDescription
parameters.control_conditionstringYesControl condition extracted from the input video:
posebodyface: facial expressions + body motion
posebody: body motion only
depth: composition and motion contours
scribble: line-art structure
parameters.strengthfloatNo1.0Control strength, range [0.0, 1.0]; larger values stay closer to the motion and composition of the original video

video_edit (Video Edit / inpainting)#

input-specific fields:
ParameterTypeRequiredDefaultDescription
input.video_urlstringYesInput video URL; limits are described in "Common Input Material Limits"
input.ref_images_urlarray[string]No1 image only, usable as a subject or background image, to replace the content of the edited region
input.mask_image_urlstringNoMask image URL, mutually exclusive with mask_video_url (this parameter is preferred). White [255,255,255] = edit region, black [0,0,0] = preserved region. Its resolution must be exactly identical to the input video
input.mask_frame_idintegerNo1Effective when mask_image_url is non-empty; the frame ID where the mask target is located. Range [1, max_frame_id], where max_frame_id = frame rate × duration + 1 (e.g. 16FPS × 5s → 81)
input.mask_video_urlstringNoMask video URL, mutually exclusive with mask_image_url. Its format / frame rate / resolution / length must be fully identical to the input video; black/white semantics are the same as above
parameters-specific fields:
ParameterTypeRequiredDefaultDescription
parameters.control_conditionstringNoEmpty (no extraction)Options posebodyface, depth; posebodyface suits scenes where the subject's face is large in frame with clear features
parameters.mask_typestringNotrackingEffective when mask_image_url is non-empty: tracking = the edit region follows the target's motion trajectory; fixed = the edit region stays fixed
parameters.expand_ratiofloatNo0.05Effective when mask_type=tracking; outward expansion ratio of the mask, range [0.0, 1.0]; the default value is recommended
parameters.expand_modestringNohullEffective when mask_type=tracking; shape of the mask region: hull (polygon) / bbox (rectangular bounding box) / original (fits the original shape)
parameters.sizestringNo1280*720Same five options as for image_reference

video_extension (Video Extension)#

input-specific fields (first/last frame images and first/last video clips are all optional; combine as needed, but at least one must be provided):
ParameterTypeRequiredDefaultDescription
input.first_frame_urlstringNoFirst frame image URL; extends forward from this frame
input.last_frame_urlstringNoLast frame image URL; converges toward this frame
input.first_clip_urlstringNoFirst video clip URL, length ≤ 3 seconds (only the first 3 seconds are used if longer); when used together with a last clip, the two clips' combined duration must be ≤ 3 seconds, and matching frame rates are recommended
input.last_clip_urlstringNoLast video clip URL; same limits as first_clip_url
input.video_urlstringNoReference video, used to extract motion features to guide generation, in combination with first/last frames or clips. Length ≤ 5 seconds; frame rate ≥ 16FPS and consistent with the first/last clips; resolution consistent with the first/last frames/clips
parameters-specific fields:
ParameterTypeRequiredDefaultDescription
parameters.control_conditionstringRequired when video_url is providedEmpty (no extraction)Options posebodyface, depth
Reminder: the output has a fixed total duration of 5 seconds, which includes the provided first/last clips themselves — it is not 5 seconds appended to the original video.

video_outpainting (Video Outpainting)#

input-specific fields:
ParameterTypeRequiredDefaultDescription
input.video_urlstringYesInput video URL; limits are described in "Common Input Material Limits"
parameters-specific fields (all four direction ratios are relative to the centered frame; 1.0 = no expansion):
ParameterTypeRequiredDefaultDescription
parameters.top_scalefloatNo1.0Upward expansion ratio, range [1.0, 2.0]
parameters.bottom_scalefloatNo1.0Downward expansion ratio, range [1.0, 2.0]
parameters.left_scalefloatNo1.0Leftward expansion ratio, range [1.0, 2.0]
parameters.right_scalefloatNo1.0Rightward expansion ratio, range [1.0, 2.0]

Common Input Material Limits#

MaterialLimits
Images (reference / mask / first-last frames)Format JPG, JPEG, PNG, BMP, TIFF, WEBP; width and height within [360, 2000] pixels; ≤ 10MB; URL must not contain Chinese characters. Mask image resolution must be exactly identical to the input video
Videos (video_url, etc.)Format MP4; frame rate ≥ 16FPS; ≤ 50MB; length ≤ 5 seconds (only the first 5 seconds are used if longer; video_extension clips ≤ 3 seconds); URL must not contain Chinese characters
URL formPublicly accessible HTTP/HTTPS address
Number of reference imagesimage_reference: 1–3 images (at most 1 bg); video_repainting / video_edit: 1 image only

Request Example: Multi-Image Reference to video (image_reference)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wanx2.1-vace-plus",
    "input": {
      "function": "image_reference",
      "prompt": "In the video, a girl gracefully walks out from the depths of an ancient forest shrouded in morning mist, her steps light",
      "ref_images_url": [
        "http://wanx.alicdn.com/material/20250318/image_reference_2_5_16.png",
        "http://wanx.alicdn.com/material/20250318/image_reference_1_5_16.png"
      ]
    },
    "parameters": {
      "prompt_extend": true,
      "obj_or_bg": ["obj", "bg"],
      "size": "1280*720"
    }
  }'

Request Example: Video Repainting (video_repainting)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wanx2.1-vace-plus",
    "input": {
      "function": "video_repainting",
      "prompt": "The video shows a black steampunk-style car driven by a gentleman, the vehicle adorned with gears and copper pipes",
      "video_url": "http://wanx.alicdn.com/material/20250318/video_repainting_1.mp4"
    },
    "parameters": {
      "prompt_extend": false,
      "control_condition": "depth"
    }
  }'

Request Example: Video Extension (video_extension)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wanx2.1-vace-plus",
    "input": {
      "function": "video_extension",
      "prompt": "A dog wearing sunglasses skateboarding down the street, 3D cartoon style",
      "first_clip_url": "http://wanx.alicdn.com/material/20250318/video_extension_1.mp4"
    },
    "parameters": {
      "prompt_extend": false
    }
  }'

Response Example#

{
  "output": {
    "task_status": "PENDING",
    "task_id": "0385dc79-5ff8-4d82-bcb6-xxxxxx"
  },
  "request_id": "4909100c-7b5a-9f92-bfe5-xxxxxx"
}

Response Fields#

FieldTypeDescription
output.task_idstringTask ID for querying the result, valid for 24 hours
output.task_statusstringTask status, PENDING on successful creation
request_idstringUnique request identifier, used for tracing and troubleshooting
code / messagestringReturned only when creation fails; error code and details (e.g. InvalidParameter: size mismatch, obj_or_bg length inconsistent with ref_images_url, etc.)

Billing#

Billed by the number of seconds of output video (a qualitative description; unit prices are subject to the platform's pricing page). After a task succeeds, the usage field in the query result is the billing basis; output is fixed at the 720P tier, up to 5 seconds. Failed tasks incur no charges.

Query Task#

Poll the task status using the task_id returned when creating the task; after success, retrieve the video from output.video_url (valid for 24 hours).
GET https://platform.dataeyes.ai/ali/api/v1/tasks/{task_id}
See 09-Query-Task for details.
Previous
07-Digital-Human-Video
Next
09-Query-Task