DataEyesAI
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
  1. Wan Video Generation
  • Getting Started
    • Overview
    • Console (Getting Started)
    • API Key
    • Base URL
  • Developer Tool Integration
    • OpenClaw
    • Claude Code
    • Codex
    • Gemini CLI
    • Grok CLI
    • Other Tools
  • AI Models API
    • OpenAI format (supports major original models)
      • Chat (Response)
        • Create Network Search
        • Create Model Response GPT-5 Enable Thinking
        • Create Function Call
        • Create Model Response
        • Create Model Response (Streaming Return)
        • Create Model Response (Control Thinking Length)
      • ChatGPT Interface
        • Audio
          • Audio to text gpt-4o-transcribe
          • GPT-4o-audio
          • Audio to text whisper-1
          • Audio to text gpt-4o-transcribe
          • Create voice gpt-4o-mini-tts
        • Chat
          • Create chat-based image recognition (non-streaming)
          • Create chat-based image recognition (streaming)
          • Create chat-based image recognition (streaming) best64
          • Official N test
          • Create structured output
          • Control the effort level of the inference model
          • Create chat function call
          • deepseek-ocr recognition
          • Create chat completion (non-stream)
        • Completions
          • ChatGPT automatic completion
          • Create completion
      • Image
        • Edit image
        • Create chat completion (streaming)
        • Create chat completion (qwen-mt-turbo)
        • Create chat completion with deepseek v3.1 level of reasoning (streaming)
      • Audio
        • Speech recognition
        • Speech synthesis
        • Official Function Calling invocation
        • Create chat-generated images (non-streaming)
      • Embedding
        • Text embeddings
    • Anthropic format
      • Chat
      • Chat(prompt cache)
      • Streaming response
      • Chat (deep reasoning)
      • Tool invocation (function call)
      • Analyze image
    • Google Gemini interface
      • Native format
        • Text-to-image + control over aspect ratio + clarity
        • Generate image
        • Text generation
        • Text generation - stream
        • Text generation + reasoning - stream
        • Image generation
        • Formatted output
        • Function call
        • Document understanding
        • URL context [native format]
        • Code execution
        • Video understanding
        • URL context
        • Video understanding - url [native format]
        • Imagen 4
        • Audio understanding
        • Embeddings
        • Chat
        • Edit image
      • Image-to-image Base64 request method
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
        • Image editing
        • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image generation( gemini-2.5-flash-image)
        • Image generation gemini-2.5-flash-image, controlling aspect ratio.
        • Image understanding
      • Image-to-image URL request returns URL request format OpenAI
        • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image understanding
    • NanoBanana
      • OpenAI request
        • Edit image
        • OpenAI image format
      • Gemini request
        • Generate image
        • Edit image
    • Midjourney format
      • Midjourney API Reference
      • Task query interface
      • Upload image
      • Get seed (Seed)
      • Submit Imagine task
      • Query tasks based on ID list
      • FaceSwap
      • Execute Action operation
      • /mj/submit/blend
      • Submit Describe task
      • Submit Modal
      • Refresh link
      • Edit image
      • Query task status by task ID
      • Get the seed of the task image
    • Doubao - Painting
      • doubao-seededit-3-0-i2i-250628
      • doubao-seedream-4-0-250828 - text-to-image
      • doubao-seedream-4-0-250828 - image-to-image
      • doubao-seedream-4-0-250828 - multi-image generation
    • Rerank Reordering Model
      • Rerank
    • Video Model
      • Grok Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-Video-Extension
      • Seedance Video Generation
        • 00-Overview
        • 01-Create-Video-Generation-Task
        • 02-Query-Video-Generation-Task
        • 03-Query-Video-Generation-Task-List
        • 04-Cancel-or-Delete-Task
        • Seedance Private Asset Library API Documentation
      • MiniMax-H3 Video Generation
        • 00-Overview
        • 01-Create-Video-Generation
        • 02-Create-Video-Regeneration
        • 03-Create-H3-Context-IR
        • 04-Query-Task
        • 05-List-Tasks
        • 06-Cancel-or-Delete-Task
      • Hailuo Video Generation
        • 00-Overview
        • 01-Text-to-Video-T2V
        • 02-Image-to-Video-I2V
        • 03-First-Last-Frame-FL2V
        • 04-Subject-Reference-S2V
        • 05-Query-Task-Status
        • 06-Video-Download
        • 99-Appendix-Camera-Movement-and-Webhooks
      • Jimeng Video Generation
        • 00-Overview
        • 01-3.0-Pro-Video-Generation
        • 02-720P-Text-to-Video
        • 03-720P-Image-to-Video-First-Frame
        • 04-720P-Image-to-Video-Start-End-Frame
        • 05-720P-Image-to-Video-Camera
        • 06-1080P-Text-to-Video
        • 07-1080P-Image-to-Video-First-Frame
        • 08-1080P-Image-to-Video-Start-End-Frame
        • 09-Error-Codes
      • Kling AI Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Omni-Video
        • 04-Multi-Image-to-Video
        • 05-Motion-Control
        • 06-Multi-Elements
        • 07-Video-Extension
        • 08-Lip-Sync
        • 09-Avatar
        • 10-Text-to-Audio
        • 11-Video-to-Audio
        • 12-TTS
        • 13-Custom-Voices
        • 14-Image-Recognition
        • 15-Element-Management
        • 16-Video-Effects
      • Vidu Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Start-End-Frame
        • 05-Multi-Frame
        • 06-Scene-Template
        • 07-Template-Story
        • 08-Query-Tasks
      • HappyHorse
        • HappyHorse Text-to-Video
        • HappyHorse Image-to-Video (First Frame)
        • HappyHorse Reference-to-Video
        • HappyHorse Video Editing
      • Wan Video Generation
        • 00-Overview.md
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-First-Last-Frame-to-Video
        • 06-Motion-Transfer-and-Character-Swap
        • 07-Digital-Human-Video
        • 08-VACE-Video-Editing
        • 09-Query-Task
    • Audio API
      • Audio API
      • Gemini TTS API
      • Google DeepMind Lyria API
      • Elevenlabs Speech to Text API Reference
      • Text-to-Music Suno
        • Task Submission
          • Generate Song (Inspiration Mode)
          • Generate Song (Custom Mode)
          • Generate Song (Continuation Mode)
          • Generate Song (Singer Style)
          • Generate Song (Secondary Creation from Uploaded Song)
          • Generate Song (Song Stitching)
          • Generate Lyrics
          • Song Stitching
        • Query Interface
          • Batch Retrieve Tasks
          • Query Single Task
  • Search / Reader Product
    • Web Reader API​​
      • Web Reader API
      • Web Reader API(HK)
    • Web Search API​​
      • Modal Card API
        • Weather
          • All City ID
          • Weather Query API
      • Web Search API
      • Video Search api
      • Trending Search API
    • Document OCR Parsing API
      • fiel upload
      • URL Parsing
  • Advanced & System API
    • Data Updates
    • System interface
      • API Key & Quota Query API
      • API Key Management API
    • API Reference​​
      • Error Codes
      • HTTP Notes
    • List models
      • Models
  1. Wan Video Generation

07-Digital-Human-Video

Digital Human Video#

Doc version:v1.0.0 | Last updated:2026-07-22
This platform fully supports the official Tongyi Wanxiang (Wan) video generation API. Requests and responses are transparently proxied; parameter semantics are identical to the official API.
Based on a single portrait image + a vocal audio clip, generate a talking, singing, or performing video with lip movements, facial expressions, and body motions synchronized to the audio. Supports real people (portrait / half-body / full-body) as well as cartoon characters.
This document covers two endpoints:
EndpointModelInvocation
Digital Human Video generationwan2.2-s2vAsynchronous (create task + poll for results)
Face Detectionwan2.2-s2v-detectSynchronous (returns result immediately)
Recommended workflow: Call the Face Detection endpoint first to confirm the image is compliant (check_pass: true), then submit the Digital Human Video generation task. This avoids task failures caused by non-compliant images.

1. Digital Human Video Generation (wan2.2-s2v)#

Create Task#

POST https://platform.dataeyes.ai/ali/api/v1/services/aigc/image2video/video-synthesis

Request Headers#

ParameterTypeRequiredDefaultDescription
Content-TypestringYesapplication/jsonData exchange format
AuthorizationstringYesAuthentication, Bearer {API_KEY}
The X-DashScope-Async: enable header required by the official API is added automatically by the platform — you do not need to include it; including it is also compatible.

Request Body#

ParameterTypeRequiredDefaultDescription
modelstringYesFixed to wan2.2-s2v
input.image_urlstringYesPortrait image URL; must be a publicly accessible HTTP/HTTPS link. See "Image and Audio Requirements" below
input.audio_urlstringYesVocal audio URL; must be a publicly accessible HTTP/HTTPS link. See "Image and Audio Requirements" below
parameters.resolutionstringNo480POutput resolution tier, options 480P, 720P. Output keeps the same aspect ratio as the input image, with total pixels adjusted to around the selected tier (480P ≈ 310K pixels, 720P ≈ 920K pixels; e.g. a 4:5 input image at 480P may output 480×600)
parameters.stylestringNoLip-sync scenario, e.g. speech (talking). The official documentation describes three scenarios — talking, singing, and performing — but the enum values for singing/performing are not publicly documented; verify through testing

Video Duration Rules#

Generated video duration = input audio duration; it is not specified via a parameter:
The audio duration must be less than 20 seconds, otherwise the task will fail with an error;
For longer videos, split the audio into segments of less than 20 seconds each, generate them separately, then concatenate the results.

Image and Audio Requirements#

InputRequirements
ImageFormat jpg/jpeg/png/bmp/webp; both width and height within [400, 7000] pixels; a clear, front-facing, single-person portrait is recommended (you can pre-check with the Face Detection endpoint)
AudioFormat wav/mp3; file smaller than 15MB, duration less than 20 seconds; must contain clear and loud vocals — removing ambient noise and background music is recommended

Request Example (480P)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/image2video/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.2-s2v",
    "input": {
      "image_url": "https://example.com/portrait.jpg",
      "audio_url": "https://example.com/speech.mp3"
    },
    "parameters": {
      "resolution": "480P",
      "style": "speech"
    }
  }'

Request Example (720P)#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/image2video/video-synthesis' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.2-s2v",
    "input": {
      "image_url": "https://example.com/portrait.jpg",
      "audio_url": "https://example.com/sing.wav"
    },
    "parameters": {
      "resolution": "720P",
      "style": "speech"
    }
  }'

Create Response Example#

{
  "output": {
    "task_status": "PENDING",
    "task_id": "0385dc79-5ff8-4d82-bcb6-xxxxxx"
  },
  "request_id": "4909100c-7b5a-9f92-bfe5-xxxxxx"
}

Usage and Billing#

After the task succeeds, in the usage object returned by the query endpoint:
usage.duration: duration of the generated video in seconds, equal to the input audio duration — this is the billing basis;
usage.SR: output resolution tier (480 or 720); unit prices differ by tier, with 720P priced higher than 480P.
Digital Human Video is billed by the number of seconds of successfully generated video and the resolution tier; inputs are not billed; failed or errored tasks are not billed. Generation typically takes about 5–10 minutes; polling the query endpoint at roughly 15-second intervals is recommended.

2. Face Detection (wan2.2-s2v-detect)#

Pre-checks whether the input image meets the portrait requirements for Digital Human Video generation (e.g. clarity, single person, front-facing), returning whether it passes and whether a human figure is detected.
POST https://platform.dataeyes.ai/ali/api/v1/services/aigc/image2video/face-detect
This is a synchronous endpoint: the detection result is returned immediately in the response, no task_id is produced, and no polling is needed.

Request Headers#

ParameterTypeRequiredDefaultDescription
Content-TypestringYesapplication/jsonData exchange format
AuthorizationstringYesAuthentication, Bearer {API_KEY}

Request Body#

ParameterTypeRequiredDefaultDescription
modelstringYesFixed to wan2.2-s2v-detect
input.image_urlstringYesURL of the image to check; must be a publicly accessible HTTP/HTTPS link. Format jpg/jpeg/png/bmp/webp; both width and height within [400, 7000] pixels

Request Example#

curl --request POST \
  --url 'https://platform.dataeyes.ai/ali/api/v1/services/aigc/image2video/face-detect' \
  --header 'Authorization: Bearer <token>' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "wan2.2-s2v-detect",
    "input": {
      "image_url": "https://example.com/portrait.jpg"
    }
  }'

Response Example#

Check passed:
{
  "output": {
    "check_pass": true,
    "humanoid": true
  },
  "usage": {
    "image_count": 1
  },
  "request_id": "c56f62df-xxxx-xxxx-xxxx-xxxxxx"
}
Check failed:
{
  "output": {
    "check_pass": false,
    "humanoid": false,
    "code": "xxx",
    "message": "xxx"
  },
  "usage": {
    "image_count": 1
  },
  "request_id": "c56f62df-xxxx-xxxx-xxxx-xxxxxx"
}

Response Fields#

FieldTypeDescription
output.check_passbooleanDetection result: true = passed, the image can be used for Digital Human Video generation; false = failed
output.humanoidbooleanWhether a human figure was detected
output.code / output.messagestringReturned only when the check fails, giving the reason; absent when the check passes
usage.image_countintegerNumber of images checked in this request, fixed at 1 — this is the billing basis
request_idstringUnique ID of this request
Face Detection is billed by the number of images checked: a successful request is billed regardless of whether the check passes; failed requests are not billed.

Query Task#

Digital Human Video generation (wan2.2-s2v) is an asynchronous task; after creation, poll for status and results using task_id (Face Detection is synchronous and requires no querying):
GET https://platform.dataeyes.ai/ali/api/v1/tasks/{task_id}
See 09-Query-Task for details.
Previous
06-Motion-Transfer-and-Character-Swap
Next
08-VACE-Video-Editing