DataEyesAI
Official SiteConsoleDocs HomeGetting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs HomeGetting StartedDeveloper ToolsAI Models API
  1. Wan Video Generation
  • OpenAI format (supports major original models)
    • Chat (Response)
      • Create Network Search
      • Create Model Response GPT-5 Enable Thinking
      • Create Function Call
      • Create Model Response
      • Create Model Response (Streaming Return)
      • Create Model Response (Control Thinking Length)
    • ChatGPT Interface
      • Audio
        • Audio to text gpt-4o-transcribe
        • GPT-4o-audio
        • Audio to text whisper-1
        • Audio to text gpt-4o-transcribe
        • Create voice gpt-4o-mini-tts
      • Chat
        • Create chat-based image recognition (non-streaming)
        • Create chat-based image recognition (streaming)
        • Create chat-based image recognition (streaming) best64
        • Official N test
        • Create structured output
        • Control the effort level of the inference model
        • Create chat function call
        • deepseek-ocr recognition
        • Create chat completion (non-stream)
      • Completions
        • ChatGPT automatic completion
        • Create completion
    • Image
      • Edit image
      • Create chat completion (streaming)
      • Create chat completion (qwen-mt-turbo)
      • Create chat completion with deepseek v3.1 level of reasoning (streaming)
    • Audio
      • Speech recognition
      • Speech synthesis
      • Official Function Calling invocation
      • Create chat-generated images (non-streaming)
    • Embedding
      • Text embeddings
  • Anthropic format
    • Chat
    • Chat(prompt cache)
    • Streaming response
    • Chat (deep reasoning)
    • Tool invocation (function call)
    • Analyze image
  • Google Gemini interface
    • Native format
      • Text-to-image + control over aspect ratio + clarity
      • Generate image
      • Text generation
      • Text generation - stream
      • Text generation + reasoning - stream
      • Image generation
      • Formatted output
      • Function call
      • Document understanding
      • URL context [native format]
      • Code execution
      • Video understanding
      • URL context
      • Video understanding - url [native format]
      • Imagen 4
      • Audio understanding
      • Embeddings
      • Chat
      • Edit image
    • Image-to-image Base64 request method
      • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
      • Image editing
      • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Image generation( gemini-2.5-flash-image)
      • Image generation gemini-2.5-flash-image, controlling aspect ratio.
      • Image understanding
    • Image-to-image URL request returns URL request format OpenAI
      • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Image understanding
  • NanoBanana
    • OpenAI request
      • Edit image
      • OpenAI image format
    • Gemini request
      • Generate image
      • Edit image
  • Midjourney format
    • Midjourney API Reference
    • Task query interface
    • Upload image
    • Get seed (Seed)
    • Submit Imagine task
    • Query tasks based on ID list
    • FaceSwap
    • Execute Action operation
    • /mj/submit/blend
    • Submit Describe task
    • Submit Modal
    • Refresh link
    • Edit image
    • Query task status by task ID
    • Get the seed of the task image
  • Doubao - Painting
    • doubao-seededit-3-0-i2i-250628
    • doubao-seedream-4-0-250828 - text-to-image
    • doubao-seedream-4-0-250828 - image-to-image
    • doubao-seedream-4-0-250828 - multi-image generation
  • Rerank Reordering Model
    • Rerank
  • Video Model
    • Grok Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Video-Editing
      • 05-Video-Extension
    • Seedance Video Generation
      • 00-Overview
      • 01-Create-Video-Generation-Task
      • 02-Query-Video-Generation-Task
      • 03-Query-Video-Generation-Task-List
      • 04-Cancel-or-Delete-Task
      • Seedance Private Asset Library API Documentation
    • MiniMax-H3 Video Generation
      • 00-Overview
      • 01-Create-Video-Generation
      • 02-Create-Video-Regeneration
      • 03-Create-H3-Context-IR
      • 04-Query-Task
      • 05-List-Tasks
      • 06-Cancel-or-Delete-Task
    • Hailuo Video Generation
      • 00-Overview
      • 01-Text-to-Video-T2V
      • 02-Image-to-Video-I2V
      • 03-First-Last-Frame-FL2V
      • 04-Subject-Reference-S2V
      • 05-Query-Task-Status
      • 06-Video-Download
      • 99-Appendix-Camera-Movement-and-Webhooks
    • Jimeng Video Generation
      • 00-Overview
      • 01-3.0-Pro-Video-Generation
      • 02-720P-Text-to-Video
      • 03-720P-Image-to-Video-First-Frame
      • 04-720P-Image-to-Video-Start-End-Frame
      • 05-720P-Image-to-Video-Camera
      • 06-1080P-Text-to-Video
      • 07-1080P-Image-to-Video-First-Frame
      • 08-1080P-Image-to-Video-Start-End-Frame
      • 09-Error-Codes
    • Kling AI Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Omni-Video
      • 04-Multi-Image-to-Video
      • 05-Motion-Control
      • 06-Multi-Elements
      • 07-Video-Extension
      • 08-Lip-Sync
      • 09-Avatar
      • 10-Text-to-Audio
      • 11-Video-to-Audio
      • 12-TTS
      • 13-Custom-Voices
      • 14-Image-Recognition
      • 15-Element-Management
      • 16-Video-Effects
    • Vidu Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Start-End-Frame
      • 05-Multi-Frame
      • 06-Scene-Template
      • 07-Template-Story
      • 08-Query-Tasks
    • HappyHorse
      • HappyHorse Text-to-Video
      • HappyHorse Image-to-Video (First Frame)
      • HappyHorse Reference-to-Video
      • HappyHorse Video Editing
    • Wan Video Generation
      • 00-Overview.md
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Video-Editing
      • 05-First-Last-Frame-to-Video
      • 06-Motion-Transfer-and-Character-Swap
      • 07-Digital-Human-Video
      • 08-VACE-Video-Editing
      • 09-Query-Task
  • Audio API
    • Audio API
    • Gemini TTS API
    • Google DeepMind Lyria API
    • Elevenlabs Speech to Text API Reference
    • Text-to-Music Suno
      • Task Submission
        • Generate Song (Inspiration Mode)
        • Generate Song (Custom Mode)
        • Generate Song (Continuation Mode)
        • Generate Song (Singer Style)
        • Generate Song (Secondary Creation from Uploaded Song)
        • Generate Song (Song Stitching)
        • Generate Lyrics
        • Song Stitching
      • Query Interface
        • Batch Retrieve Tasks
        • Query Single Task
  1. Wan Video Generation

03-Reference-to-Video

Reference-to-Video#

Doc version:v1.0.0 | Last updated:2026-07-22
This platform fully supports the official Tongyi Wanxiang (Wan) video generation API. Requests and responses are transparently proxied; parameter semantics are identical to the official API.
Generate videos based on the subject (person, character, object, etc.) in reference images/videos, and use reference syntax in the prompt to specify the role of each reference asset. Suitable for serialized creation scenarios that require character consistency.
Reference-to-video involves two protocols with different request body structures — be careful to distinguish them:
ModelProtocolReference asset fieldPrompt reference syntaxResolution parameterAudioMulti-shot
wan2.7-r2v, wan2.7-r2v-2026-06-12Newinput.media (array of objects, type + url)Chinese "图1 / 视频1", English "Image 1 / Video 1"; images and videos are counted separatelyparameters.resolution (720P/1080P)Generates videos with audio by default, no setting needed; supports reference_voice voice-tone referenceshot_type not supported; achieve multi-shot by writing a storyboard script in the prompt
wan2.6-r2vLegacyinput.reference_urls (array of strings)character1, character2 (images and videos are counted together in array order)parameters.size (width*height, e.g. 1280*720)audio parameter not supported (only the flash variant supports it)parameters.shot_type (single/multi)
wan2.6-r2v-flashLegacySame as wan2.6-r2vSame as aboveSame as aboveSupports audio (true = with audio / false = silent)Same as above
⚠️ Important: the reference asset field for wan2.6-r2v is reference_urls (an array, mixing images and videos). The legacy field reference_video_urls is deprecated, and made-up field names (such as ref_imgs) are rejected outright by the official API with the error please provide reference_video_urls or reference_urls.

Prompt Reference Syntax#

Reference assets are referenced by number in the prompt; the number matches the asset's position in the array:
wan2.7-r2v (images and videos counted separately)
In Chinese, use "图1、图2" to refer to reference images and "视频1、视频2" to refer to reference videos; in English, write Image 1, Video 1 (space between word and number, first letter capitalized).
The 1st reference_image is "Image 1", and the 1st reference_video is "Video 1"; both can be present at the same time.
If there is only one image / one video, you can simply write "the reference image" / "the reference video".
Both phrasings are supported: "[Image 1] playing inside [Image 2]" and "the cat from [Image 1] playing in the room from [Image 2]".
When the reference image is a multi-panel grid (storyboard), it is recommended to describe it as a storyboard script, and pass only one multi-panel image per request.
wan2.6-r2v (counted together)
Use character1, character2 to reference the reference characters: the 1st URL in the reference_urls array is character1, the 2nd is character2, and so on (no distinction between images and videos).
Each reference asset should contain only a single character.

Create Task#

POST https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis

Request Headers#

ParameterTypeRequiredDefaultDescription
Content-TypestringYesapplication/jsonData exchange format
AuthorizationstringYesAuthentication, Bearer {API_KEY}
The X-DashScope-Async: enable header required by the official API is automatically added by the platform — no need to include it in your request.

Request Body — wan2.7-r2v (new protocol)#

ParameterTypeRequiredDefaultDescription
modelstringYeswan2.7-r2v or wan2.7-r2v-2026-06-12
input.promptstringYesPrompt, Chinese or English, ≤5000 characters (truncated if longer)
See "Prompt Reference Syntax" above for reference syntax
input.negative_promptstringNoNegative prompt, ≤500 characters
input.mediaarray of objectYesArray of reference assets. At most 1 first-frame image; at least 1 reference image + reference video and no more than 5 in total; when an asset serves as a subject character it should contain only a single character
input.media[].typestringYesAsset type:
reference_image — reference image (subject character / scene reference)
reference_video — reference video (subject character + voice-tone reference; empty establishing shots are not recommended)
first_frame — first-frame image (can be combined with subject references for joint control)
input.media[].urlstringYesAsset location. Images support public HTTP(S) URLs or Base64 (data:{MIME_type};base64,{data}); videos support public HTTP(S) URLs
input.media[].reference_voicestringNoAudio URL specifying the voice tone of this asset's subject character (voice tone reference only, unrelated to spoken content). If the reference video contains audio and this is not specified, the video's original audio is used by default; if both are provided, reference_voice takes precedence. It is recommended that the audio language match the prompt language
parameters.resolutionstringNo1080PResolution tier, options 720P/1080P; affects cost
parameters.ratiostringNo16:9Aspect ratio, options 16:9/9:16/1:1/4:3/3:4
Automatically ignored when a first-frame image is provided; the output uses an approximate ratio based on the first frame's aspect ratio
parameters.durationintegerNo5Video duration in seconds; affects cost
Integer in 2–10 when reference assets include a video; integer in 2–15 when they do not
parameters.prompt_extendbooleanNotrueIntelligent prompt rewriting; noticeable improvement for short prompts, but increases processing time
parameters.watermarkbooleanNofalse"AI-generated" watermark in the bottom-right corner
parameters.seedintegerNoRandomRange [0, 2147483647]; identical seeds still do not guarantee identical results

Request Body — wan2.6-r2v (legacy protocol)#

ParameterTypeRequiredDefaultDescription
modelstringYeswan2.6-r2v or wan2.6-r2v-flash
input.promptstringYesPrompt, ≤1500 characters
Use character1, character2 to reference the reference characters (counted together in array order)
input.negative_promptstringNoNegative prompt, ≤500 characters
input.reference_urlsarray of stringYesArray of reference file URLs, mixing images and videos: 0–5 images, 0–3 videos, images + videos ≤ 5 in total; affects cost
⚠️ The legacy field reference_video_urls is deprecated — do not use it
parameters.sizestringNo1920*1080Output resolution; must be specified as exact pixels "width*height" (not 1:1 or 720P); affects cost
720P tier: 1280*720(16:9), 720*1280(9:16), 960*960(1:1), 1088*832(4:3), 832*1088(3:4)
1080P tier: 1920*1080(16:9), 1080*1920(9:16), 1440*1440(1:1), 1632*1248(4:3), 1248*1632(3:4)
parameters.durationintegerNo5Video duration in seconds, integer in 2–10; affects cost
parameters.shot_typestringNosingleShot type: single single shot / multi multi-shot; takes precedence over descriptions in the prompt
parameters.audiobooleanNotrueOnly supported by wan2.6-r2v-flash. true with audio / false silent (must be explicitly set to false to generate a silent video); affects cost
parameters.watermarkbooleanNofalse"AI-generated" watermark in the bottom-right corner
parameters.seedintegerNoRandomRange [0, 2147483647]

Reference Asset Constraints#

AssetCountFormatDimensions / DurationSize
Reference imagewan2.7: images + videos ≤ 5 in total; wan2.6: 0–5 images and images + videos ≤ 5 in totalJPEG, JPG, PNG (alpha channel not supported), BMP, WEBPWidth/height 240–8000 px; wan2.7 requires aspect ratio 1:8–8:1≤ 20MB
Reference videowan2.7: images + videos ≤ 5 in total; wan2.6: 0–3 videosMP4, MOVDuration 1–30 seconds; wan2.7 requires width/height 240–4096 px and aspect ratio 1:8–8:1≤ 100MB
First-frame image (wan2.7 first_frame only)At most 1Same as reference imageSame as reference image≤ 20MB
Reference audio (wan2.7 reference_voice only)1 per assetWAV, MP3Duration 1–10 seconds≤ 15MB

Billing Notes#

Billing is based on usage.duration in the query result: duration = input reference video duration + output video duration. That is, when a reference video is provided, its duration also counts toward the billed duration (e.g. a 5-second input reference video + 10-second output is billed as 15 seconds). With image-only reference input, input_video_duration is 0.
Parameters that affect cost: resolution (wan2.7's resolution / wan2.6's size; 1080P costs more than 720P), duration, reference assets (wan2.6's reference_urls), and audio (wan2.6-r2v-flash only; with-audio and silent are priced differently).

Request Example — wan2.7-r2v#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.7-r2v",
    "input": {
      "prompt": "The woman from Image 1 strolling gracefully through a garden, her dress swaying gently in the breeze",
      "media": [
        {
          "type": "reference_image",
          "url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260424/mvzfud/hh-v2v-girl.jpg"
        }
      ]
    },
    "parameters": {
      "resolution": "1080P",
      "duration": 5
    }
  }'

Request Example — wan2.6-r2v#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.6-r2v",
    "input": {
      "prompt": "The person from character1 strolling along a path covered with autumn leaves",
      "reference_urls": [
        "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260424/mvzfud/hh-v2v-girl.jpg"
      ]
    },
    "parameters": {
      "size": "1280*720",
      "duration": 5
    }
  }'

Response Example#

Success:
{
  "output": {
    "task_status": "PENDING",
    "task_id": "0385dc79-5ff8-4d82-bcb6-xxxxxx"
  },
  "request_id": "4909100c-7b5a-9f92-bfe5-xxxxxx"
}
Failed:
{
  "code": "InvalidParameter",
  "message": "please provide reference_video_urls or reference_urls",
  "request_id": "7438d53d-xxxx"
}

Response Fields#

FieldTypeDescription
output.task_idstringTask ID, used for polling; valid for 24 hours
output.task_statusstringTask status, PENDING on successful creation
request_idstringUnique request identifier
code / messagestringError code and details, returned only when creation fails

Query Task#

Poll the task status with the task_id returned by task creation; once successful, get the video URL from output.video_url:
GET https://platform.dataeyes.ai/ali/api/v1/tasks/{task_id}
See 09-Query-Task for details.
Previous
02-Image-to-Video
Next
04-Video-Editing