DataEyesAI
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
  1. Wan Video Generation
  • Getting Started
    • Overview
    • Console (Getting Started)
    • API Key
    • Base URL
  • Developer Tool Integration
    • OpenClaw
    • Claude Code
    • Codex
    • Gemini CLI
    • Grok CLI
    • Other Tools
  • AI Models API
    • OpenAI format (supports major original models)
      • Chat (Response)
        • Create Network Search
        • Create Model Response GPT-5 Enable Thinking
        • Create Function Call
        • Create Model Response
        • Create Model Response (Streaming Return)
        • Create Model Response (Control Thinking Length)
      • ChatGPT Interface
        • Audio
          • Audio to text gpt-4o-transcribe
          • GPT-4o-audio
          • Audio to text whisper-1
          • Audio to text gpt-4o-transcribe
          • Create voice gpt-4o-mini-tts
        • Chat
          • Create chat-based image recognition (non-streaming)
          • Create chat-based image recognition (streaming)
          • Create chat-based image recognition (streaming) best64
          • Official N test
          • Create structured output
          • Control the effort level of the inference model
          • Create chat function call
          • deepseek-ocr recognition
          • Create chat completion (non-stream)
        • Completions
          • ChatGPT automatic completion
          • Create completion
      • Image
        • Edit image
        • Create chat completion (streaming)
        • Create chat completion (qwen-mt-turbo)
        • Create chat completion with deepseek v3.1 level of reasoning (streaming)
      • Audio
        • Speech recognition
        • Speech synthesis
        • Official Function Calling invocation
        • Create chat-generated images (non-streaming)
      • Embedding
        • Text embeddings
    • Anthropic format
      • Chat
      • Chat(prompt cache)
      • Streaming response
      • Chat (deep reasoning)
      • Tool invocation (function call)
      • Analyze image
    • Google Gemini interface
      • Native format
        • Text-to-image + control over aspect ratio + clarity
        • Generate image
        • Text generation
        • Text generation - stream
        • Text generation + reasoning - stream
        • Image generation
        • Formatted output
        • Function call
        • Document understanding
        • URL context [native format]
        • Code execution
        • Video understanding
        • URL context
        • Video understanding - url [native format]
        • Imagen 4
        • Audio understanding
        • Embeddings
        • Chat
        • Edit image
      • Image-to-image Base64 request method
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
        • Image editing
        • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image generation( gemini-2.5-flash-image)
        • Image generation gemini-2.5-flash-image, controlling aspect ratio.
        • Image understanding
      • Image-to-image URL request returns URL request format OpenAI
        • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image understanding
    • NanoBanana
      • OpenAI request
        • Edit image
        • OpenAI image format
      • Gemini request
        • Generate image
        • Edit image
    • Midjourney format
      • Midjourney API Reference
      • Task query interface
      • Upload image
      • Get seed (Seed)
      • Submit Imagine task
      • Query tasks based on ID list
      • FaceSwap
      • Execute Action operation
      • /mj/submit/blend
      • Submit Describe task
      • Submit Modal
      • Refresh link
      • Edit image
      • Query task status by task ID
      • Get the seed of the task image
    • Doubao - Painting
      • doubao-seededit-3-0-i2i-250628
      • doubao-seedream-4-0-250828 - text-to-image
      • doubao-seedream-4-0-250828 - image-to-image
      • doubao-seedream-4-0-250828 - multi-image generation
    • Rerank Reordering Model
      • Rerank
    • Video Model
      • Grok Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-Video-Extension
      • Seedance Video Generation
        • 00-Overview
        • 01-Create-Video-Generation-Task
        • 02-Query-Video-Generation-Task
        • 03-Query-Video-Generation-Task-List
        • 04-Cancel-or-Delete-Task
        • Seedance Private Asset Library API Documentation
      • MiniMax-H3 Video Generation
        • 00-Overview
        • 01-Create-Video-Generation
        • 02-Create-Video-Regeneration
        • 03-Create-H3-Context-IR
        • 04-Query-Task
        • 05-List-Tasks
        • 06-Cancel-or-Delete-Task
      • Hailuo Video Generation
        • 00-Overview
        • 01-Text-to-Video-T2V
        • 02-Image-to-Video-I2V
        • 03-First-Last-Frame-FL2V
        • 04-Subject-Reference-S2V
        • 05-Query-Task-Status
        • 06-Video-Download
        • 99-Appendix-Camera-Movement-and-Webhooks
      • Jimeng Video Generation
        • 00-Overview
        • 01-3.0-Pro-Video-Generation
        • 02-720P-Text-to-Video
        • 03-720P-Image-to-Video-First-Frame
        • 04-720P-Image-to-Video-Start-End-Frame
        • 05-720P-Image-to-Video-Camera
        • 06-1080P-Text-to-Video
        • 07-1080P-Image-to-Video-First-Frame
        • 08-1080P-Image-to-Video-Start-End-Frame
        • 09-Error-Codes
      • Kling AI Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Omni-Video
        • 04-Multi-Image-to-Video
        • 05-Motion-Control
        • 06-Multi-Elements
        • 07-Video-Extension
        • 08-Lip-Sync
        • 09-Avatar
        • 10-Text-to-Audio
        • 11-Video-to-Audio
        • 12-TTS
        • 13-Custom-Voices
        • 14-Image-Recognition
        • 15-Element-Management
        • 16-Video-Effects
      • Vidu Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Start-End-Frame
        • 05-Multi-Frame
        • 06-Scene-Template
        • 07-Template-Story
        • 08-Query-Tasks
      • HappyHorse
        • HappyHorse Text-to-Video
        • HappyHorse Image-to-Video (First Frame)
        • HappyHorse Reference-to-Video
        • HappyHorse Video Editing
      • Wan Video Generation
        • 00-Overview.md
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-First-Last-Frame-to-Video
        • 06-Motion-Transfer-and-Character-Swap
        • 07-Digital-Human-Video
        • 08-VACE-Video-Editing
        • 09-Query-Task
    • Audio API
      • Audio API
      • Gemini TTS API
      • Google DeepMind Lyria API
      • Elevenlabs Speech to Text API Reference
      • Text-to-Music Suno
        • Task Submission
          • Generate Song (Inspiration Mode)
          • Generate Song (Custom Mode)
          • Generate Song (Continuation Mode)
          • Generate Song (Singer Style)
          • Generate Song (Secondary Creation from Uploaded Song)
          • Generate Song (Song Stitching)
          • Generate Lyrics
          • Song Stitching
        • Query Interface
          • Batch Retrieve Tasks
          • Query Single Task
  • Search / Reader Product
    • Web Reader API​​
      • Web Reader API
      • Web Reader API(HK)
    • Web Search API​​
      • Modal Card API
        • Weather
          • All City ID
          • Weather Query API
      • Web Search API
      • Video Search api
      • Trending Search API
    • Document OCR Parsing API
      • fiel upload
      • URL Parsing
  • Advanced & System API
    • Data Updates
    • System interface
      • API Key & Quota Query API
      • API Key Management API
    • API Reference​​
      • Error Codes
      • HTTP Notes
    • List models
      • Models
  1. Wan Video Generation

02-Image-to-Video

Image-to-Video#

Doc version:v1.0.0 | Last updated:2026-07-22
This platform fully supports the official Tongyi Wanxiang (Wan) video generation API. Requests and responses are transparently proxied; parameter semantics are identical to the official API.
Generate video using an image as the first frame, combined with a text prompt. Image-to-Video has two request protocols sharing the same endpoint, distinguished by model and request body structure:
ProtocolApplicable modelsInput methodCapabilities
New (recommended)wan2.7-i2v and other wan2.7-series modelsinput.media arrayFirst-frame-to-video, first/last-frame-to-video, audio-driven generation, video continuation
Legacywan2.6-i2v, wan2.6-i2v-flash, wan2.5-i2v-preview, wan2.2-i2v-plus, wan2.2-i2v-flash, wanx2.1-i2v-turbo, wanx2.1-i2v-plusinput.img_url fieldFirst-frame-to-video only; wan2.6/wan2.5 support audio (audio_url) and auto dubbing, wan2.2/wanx2.1 support video effect templates

Create Task#

POST https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis

Request Headers#

ParameterTypeRequiredDefaultDescription
Content-TypestringYesapplication/jsonData exchange format
AuthorizationstringYesAuthentication, Bearer {API_KEY}
The X-DashScope-Async: enable header required by the official API is added automatically by the platform; you do not need to include it.

Request Body — New Protocol (wan2.7)#

ParameterTypeRequiredDefaultDescription
modelstringYesModel name, e.g. wan2.7-i2v
input.promptstringNoText prompt, Chinese or English, up to 5000 characters; excess is automatically truncated
input.negative_promptstringNoNegative prompt, up to 500 characters; excess is automatically truncated
input.mediaarrayYesList of media assets; each element is a {type, url} object
input.media[].typestringYesAsset type: first_frame (first-frame image), last_frame (last-frame image), driving_audio (driving audio), first_clip (leading video clip, for video continuation). Each type may appear at most once in the array
input.media[].urlstringYesAsset URL. Public HTTP(S) URLs are supported; images also support Base64 (data:{MIME_type};base64,{base64_data}), while audio/video do not support Base64
parameters.resolutionstringNo1080PResolution tier, 720P or 1080P; directly affects cost. Output is scaled automatically to the tier's total pixel count, keeping the aspect ratio as close to the input asset as possible
parameters.durationintegerNo5Video duration in seconds, integer in [2, 15]; directly affects cost. For video continuation, this is the final total duration (input clip + continuation), and billing is based on the total duration
parameters.prompt_extendbooleanNotrueWhether to enable intelligent prompt rewriting; noticeably improves short prompts but increases processing time
parameters.watermarkbooleanNofalseWhether to add an "AI-generated" watermark at the bottom-right corner of the video
parameters.seedintegerNoRandomRandom seed, range [0, 2147483647]; identical seeds do not guarantee identical results

Media Combination Rules#

input.media only supports the following combinations; invalid combinations return an error:
Task typeValid combinationDescription
First-frame-to-videofirst_frameGenerate a video with the image as the first frame; the model dubs automatically (background music / sound effects)
First frame + audio-drivenfirst_frame + driving_audioDrive the visuals with the provided audio (lip sync, motion beat matching)
First/last-frame-to-videofirst_frame + last_frameGenerate a transition video from the first frame to the last frame
First/last frames + audio-drivenfirst_frame + last_frame + driving_audioFirst/last-frame transition + audio-driven
Video continuationfirst_clip; first_clip + last_frameContinue after the input video clip; a last frame can optionally be specified

Request Body — Legacy Protocol (wan2.6 and Earlier)#

ParameterTypeRequiredDefaultDescription
modelstringYesModel name, e.g. wan2.6-i2v, wan2.6-i2v-flash, wan2.2-i2v-plus, wanx2.1-i2v-turbo
input.promptstringNoText prompt, Chinese or English; wan2.6/wan2.5 series up to 1500 characters, wan2.2/wanx2.1 series up to 800 characters, excess is automatically truncated. Ignored when a template effect is used
input.negative_promptstringNoNegative prompt, up to 500 characters
input.img_urlstringYesFirst-frame image URL. Public HTTP(S) URL or Base64 (data:{MIME_type};base64,...)
input.audio_urlstringNowan2.6 / wan2.5 series only. Audio URL (public URL); if provided, the video is generated with this audio, otherwise the model dubs automatically
input.templatestringNowan2.2 / wanx2.1 series only. Video effect template name (e.g. flying); when used, prompt is ignored, and effect availability depends on the model
parameters.resolutionstringNoModel-dependentResolution tier; directly affects cost. Aspect ratio is kept as close to img_url as possible. See the tier table below for per-model values
parameters.durationintegerNoModel-dependentVideo duration in seconds; directly affects cost. See the tier table below for per-model values
parameters.prompt_extendbooleanNotrueWhether to enable intelligent prompt rewriting
parameters.shot_typestringNosinglewan2.6 series only. single (single shot) / multi (multi-shot storytelling); only takes effect when prompt_extend=true, precedence: shot_type > prompt
parameters.audiobooleanNotruewan2.6-i2v-flash only. Whether to generate a video with audio; pricing differs between audio and silent output. Precedence: audio > audio_url — when audio=false, the output is silent and billed as silent even if audio_url is provided
parameters.watermarkbooleanNofalseWhether to add an "AI-generated" watermark
parameters.seedintegerNoRandomRandom seed, range [0, 2147483647]
The wan2.2 and wanx2.1 series generate silent videos by default.

Input Asset Requirements#

Images (New protocol first_frame / last_frame; legacy img_url)#

ItemRequirement
FormatJPEG, JPG, PNG (transparency not supported), BMP, WEBP
ResolutionBoth width and height within [240, 8000] pixels
Aspect ratio1:8 to 8:1
File sizewan2.7 / wan2.6 / wan2.5: up to 20MB; wan2.2 / wanx2.1: up to 10MB
Input methodPublic HTTP(S) URL, or Base64 (supported MIME types: image/jpeg, image/png, image/bmp, image/webp)

Audio (New protocol driving_audio; legacy audio_url)#

Itemwan2.7 (driving_audio)wan2.6 / wan2.5 (audio_url)
Formatwav, mp3wav, mp3
Duration2–30 seconds3–30 seconds
File size≤15MB≤15MB
Truncation ruleAudio longer than duration is trimmed to the first duration seconds; if shorter, the remainder is silent (e.g. 3s audio, 5s video → first 3s with audio, last 2s silent)Same as left
SemanticsProvided = audio-driven (lip sync, motion beat matching); omitted = auto-generated background music / sound effectsProvided = generate the video with this audio; omitted = auto dubbing
Input methodPublic HTTP(S) URL (Base64 not supported)Same as left

Video (New protocol first_clip, video continuation)#

ItemRequirement
Formatmp4, mov
Duration2–10 seconds
ResolutionBoth width and height within [240, 4096] pixels, aspect ratio 1:8 to 8:1
File size≤100MB
Continuation ruleThe continuation limit is controlled by duration: e.g. duration=15 with a 3s input clip → 12s continuation, 15s total output, billed as 15s
The output video's aspect ratio is determined by the input asset (first-frame image / leading video clip), scaled automatically to the resolution tier's total pixel count; video width and height must be multiples of 16, so there may be a slight deviation from the original image's aspect ratio.

Resolution and Duration Tiers per Model#

Modelresolution optionsresolution defaultduration values (seconds)duration default
wan2.7-i2v720P, 1080P1080PInteger [2, 15]5
wan2.6-i2v720P, 1080P1080PInteger [2, 15]5
wan2.6-i2v-flash720P, 1080P1080PInteger [2, 15]5
wan2.5-i2v-preview480P, 720P, 1080P1080P5, 105
wan2.2-i2v-flash480P, 720P, 1080P720PFixed 5, not modifiable5
wan2.2-i2v-plus480P, 1080P1080PFixed 5, not modifiable5
wanx2.1-i2v-turbo480P, 720P720P3, 4, 55
wanx2.1-i2v-plus720P720PFixed 5, not modifiable5
resolution and duration directly affect cost (billed by resolution tier and duration in seconds); wan2.6-i2v-flash has different prices for audio and silent output.

Request Examples#

Example 1: First-Frame-to-Video (wan2.7)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.7-i2v",
    "input": {
      "prompt": "The camera slowly pushes forward as light and shadow flow across the scene",
      "media": [
        {
          "type": "first_frame",
          "url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260424/mvzfud/hh-v2v-girl.jpg"
        }
      ]
    },
    "parameters": {"resolution": "1080P", "duration": 5}
  }'

Example 2: First/Last-Frame-to-Video (wan2.7)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.7-i2v",
    "input": {
      "prompt": "Petals drift down in the wind, from full bloom to withering",
      "media": [
        {
          "type": "first_frame",
          "url": "https://wanx.alicdn.com/material/20250318/first_frame.png"
        },
        {
          "type": "last_frame",
          "url": "https://wanx.alicdn.com/material/20250318/last_frame.png"
        }
      ]
    },
    "parameters": {"resolution": "720P", "duration": 5}
  }'

Example 3: First Frame + Audio-Driven (wan2.7)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.7-i2v",
    "input": {
      "prompt": "The character sways to the rhythm of the music with lively expressions",
      "media": [
        {
          "type": "first_frame",
          "url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260424/mvzfud/hh-v2v-girl.jpg"
        },
        {
          "type": "driving_audio",
          "url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250825/iaqpio/input_audio.MP3"
        }
      ]
    },
    "parameters": {"resolution": "1080P", "duration": 5}
  }'

Example 4: Legacy Protocol (wan2.6-i2v)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/video-generation/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.6-i2v",
    "input": {
      "img_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260424/mvzfud/hh-v2v-girl.jpg",
      "prompt": "The character slowly turns around, gazing into the distance"
    },
    "parameters": {"resolution": "720P", "duration": 5}
  }'

Response Example#

{
  "output": {
    "task_status": "PENDING",
    "task_id": "0385dc79-5ff8-4d82-bcb6-xxxxxx"
  },
  "request_id": "4909100c-7b5a-9f92-bfe5-xxxxxx"
}

Response Fields#

FieldTypeDescription
output.task_idstringTask ID used for polling; valid for 24 hours. Do not create duplicate tasks — just keep polling
output.task_statusstringTask status; PENDING on successful creation. See Overview for the full enum
request_idstringUnique request identifier, useful for troubleshooting
code / messagestringError code and details, returned only when creation fails (e.g. InvalidApiKey, InvalidParameter)

Query Task#

After creating a task, poll the task status with the task_id; once the status is SUCCEEDED, get the video URL from output.video_url (valid for 24 hours):
GET https://platform.dataeyes.ai/ali/api/v1/tasks/{task_id}
See 09-Query-Task for details.
Previous
01-Text-to-Video
Next
03-Reference-to-Video