DataEyesAI
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
  1. Vidu Video Generation
  • Getting Started
    • Overview
    • Console (Getting Started)
    • API Key
    • Base URL
  • Developer Tool Integration
    • OpenClaw
    • Claude Code
    • Codex
    • Gemini CLI
    • Grok CLI
    • Other Tools
  • AI Models API
    • OpenAI format (supports major original models)
      • Chat (Response)
        • Create Network Search
        • Create Model Response GPT-5 Enable Thinking
        • Create Function Call
        • Create Model Response
        • Create Model Response (Streaming Return)
        • Create Model Response (Control Thinking Length)
      • ChatGPT Interface
        • Audio
          • Audio to text gpt-4o-transcribe
          • GPT-4o-audio
          • Audio to text whisper-1
          • Audio to text gpt-4o-transcribe
          • Create voice gpt-4o-mini-tts
        • Chat
          • Create chat-based image recognition (non-streaming)
          • Create chat-based image recognition (streaming)
          • Create chat-based image recognition (streaming) best64
          • Official N test
          • Create structured output
          • Control the effort level of the inference model
          • Create chat function call
          • deepseek-ocr recognition
          • Create chat completion (non-stream)
        • Completions
          • ChatGPT automatic completion
          • Create completion
      • Image
        • Edit image
        • Create chat completion (streaming)
        • Create chat completion (qwen-mt-turbo)
        • Create chat completion with deepseek v3.1 level of reasoning (streaming)
      • Audio
        • Speech recognition
        • Speech synthesis
        • Official Function Calling invocation
        • Create chat-generated images (non-streaming)
      • Embedding
        • Text embeddings
    • Anthropic format
      • Chat
      • Chat(prompt cache)
      • Streaming response
      • Chat (deep reasoning)
      • Tool invocation (function call)
      • Analyze image
    • Google Gemini interface
      • Native format
        • Text-to-image + control over aspect ratio + clarity
        • Generate image
        • Text generation
        • Text generation - stream
        • Text generation + reasoning - stream
        • Image generation
        • Formatted output
        • Function call
        • Document understanding
        • URL context [native format]
        • Code execution
        • Video understanding
        • URL context
        • Video understanding - url [native format]
        • Imagen 4
        • Audio understanding
        • Embeddings
        • Chat
        • Edit image
      • Image-to-image Base64 request method
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
        • Image editing
        • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image generation( gemini-2.5-flash-image)
        • Image generation gemini-2.5-flash-image, controlling aspect ratio.
        • Image understanding
      • Image-to-image URL request returns URL request format OpenAI
        • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image understanding
    • NanoBanana
      • OpenAI request
        • Edit image
        • OpenAI image format
      • Gemini request
        • Generate image
        • Edit image
    • Midjourney format
      • Midjourney API Reference
      • Task query interface
      • Upload image
      • Get seed (Seed)
      • Submit Imagine task
      • Query tasks based on ID list
      • FaceSwap
      • Execute Action operation
      • /mj/submit/blend
      • Submit Describe task
      • Submit Modal
      • Refresh link
      • Edit image
      • Query task status by task ID
      • Get the seed of the task image
    • Doubao - Painting
      • doubao-seededit-3-0-i2i-250628
      • doubao-seedream-4-0-250828 - text-to-image
      • doubao-seedream-4-0-250828 - image-to-image
      • doubao-seedream-4-0-250828 - multi-image generation
    • Rerank Reordering Model
      • Rerank
    • Video Model
      • Grok Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-Video-Extension
      • Seedance Video Generation
        • 00-Overview
        • 01-Create-Video-Generation-Task
        • 02-Query-Video-Generation-Task
        • 03-Query-Video-Generation-Task-List
        • 04-Cancel-or-Delete-Task
        • Seedance Private Asset Library API Documentation
      • MiniMax-H3 Video Generation
        • 00-Overview
        • 01-Create-Video-Generation
        • 02-Create-Video-Regeneration
        • 03-Create-H3-Context-IR
        • 04-Query-Task
        • 05-List-Tasks
        • 06-Cancel-or-Delete-Task
      • Hailuo Video Generation
        • 00-Overview
        • 01-Text-to-Video-T2V
        • 02-Image-to-Video-I2V
        • 03-First-Last-Frame-FL2V
        • 04-Subject-Reference-S2V
        • 05-Query-Task-Status
        • 06-Video-Download
        • 99-Appendix-Camera-Movement-and-Webhooks
      • Jimeng Video Generation
        • 00-Overview
        • 01-3.0-Pro-Video-Generation
        • 02-720P-Text-to-Video
        • 03-720P-Image-to-Video-First-Frame
        • 04-720P-Image-to-Video-Start-End-Frame
        • 05-720P-Image-to-Video-Camera
        • 06-1080P-Text-to-Video
        • 07-1080P-Image-to-Video-First-Frame
        • 08-1080P-Image-to-Video-Start-End-Frame
        • 09-Error-Codes
      • Kling AI Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Omni-Video
        • 04-Multi-Image-to-Video
        • 05-Motion-Control
        • 06-Multi-Elements
        • 07-Video-Extension
        • 08-Lip-Sync
        • 09-Avatar
        • 10-Text-to-Audio
        • 11-Video-to-Audio
        • 12-TTS
        • 13-Custom-Voices
        • 14-Image-Recognition
        • 15-Element-Management
        • 16-Video-Effects
      • Vidu Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Start-End-Frame
        • 05-Multi-Frame
        • 06-Scene-Template
        • 07-Template-Story
        • 08-Query-Tasks
      • HappyHorse
        • HappyHorse Text-to-Video
        • HappyHorse Image-to-Video (First Frame)
        • HappyHorse Reference-to-Video
        • HappyHorse Video Editing
      • Wan Video Generation
        • 00-Overview.md
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-First-Last-Frame-to-Video
        • 06-Motion-Transfer-and-Character-Swap
        • 07-Digital-Human-Video
        • 08-VACE-Video-Editing
        • 09-Query-Task
    • Audio API
      • Audio API
      • Gemini TTS API
      • Google DeepMind Lyria API
      • Elevenlabs Speech to Text API Reference
      • Text-to-Music Suno
        • Task Submission
          • Generate Song (Inspiration Mode)
          • Generate Song (Custom Mode)
          • Generate Song (Continuation Mode)
          • Generate Song (Singer Style)
          • Generate Song (Secondary Creation from Uploaded Song)
          • Generate Song (Song Stitching)
          • Generate Lyrics
          • Song Stitching
        • Query Interface
          • Batch Retrieve Tasks
          • Query Single Task
  • Search / Reader Product
    • Web Reader API​​
      • Web Reader API
      • Web Reader API(HK)
    • Web Search API​​
      • Modal Card API
        • Weather
          • All City ID
          • Weather Query API
      • Web Search API
      • Video Search api
      • Trending Search API
    • Document OCR Parsing API
      • fiel upload
      • URL Parsing
  • Advanced & System API
    • Data Updates
    • System interface
      • API Key & Quota Query API
      • API Key Management API
    • API Reference​​
      • Error Codes
      • HTTP Notes
    • List models
      • Models
  1. Vidu Video Generation

03-Reference-to-Video

Reference-to-Video#

Document Version: v1.0.0 | Last Updated: 2026-06-11
This platform fully supports the official Vidu video generation APIs. Requests and responses are transparently proxied with identical parameter semantics.
Generate video with subject consistency from reference images/videos, with support for subject libraries. This endpoint supports two invocation modes: Subject-based invocation (via the subjects parameter) and Non-subject invocation (via the images/videos parameters).
POST https://platform.dataeyes.ai/vidu/ent/v2/reference2video

Request Parameters#

Request Headers#

HeaderRequiredDescription
Content-TypeYesapplication/json
AuthorizationYesBearer {API_KEY}

Mode 1: Subject-Based Invocation#

Pass subject information via the subjects parameter and reference them in the prompt using @subject_name.

Request Body#

ParameterSub-parameterTypeRequiredDescription
modelStringYesModel name. Options: viduq3-turbo, viduq3, viduq2-pro, viduq2, viduq1, vidu2.0.
- viduq3-turbo: Supports smart scene cutting, audio-video output, fastest generation
- viduq3: Supports smart scene cutting, audio-video output, superior multi-angle consistency
- viduq2-pro: Supports reference video, video editing, video replacement
- viduq2: Good dynamic effects, rich details
- viduq1: Clear visuals, smooth transitions, stable camera movement
- vidu2.0: Fast generation speed
auto_subjectsBoolOptionalWhether to use the smart subject library capability. Default false.
subjectsArrayYesSubject list. q3/q2/q1/2.0 models support image and text subjects only (max 7); q2-pro additionally supports video subjects (max 4 images/text, max 2 videos).
nameStringYesSubject name. Referenced in the prompt via @name.
imagesArray[String]OptionalSubject image URLs or Base64. Max 3 images. At least one of images or videos must be provided.
Supported formats: png, jpeg, jpg, webp. Base64 must include content type prefix.
videosArray[String]OptionalSubject video URLs or Base64. At least one of images or videos must be provided.
Only supported by viduq2-pro; supports 1 video of 5 seconds.
Supported formats: mp4, avi, mov.
voice_idStringOptionalVoice ID. If empty, the system automatically recommends a voice. Ineffective for q3 reference models.
server_idStringOptionalSubject ID obtained via the Create Subject API. Required when using an existing subject.
promptStringYesText prompt. Maximum 5000 characters.
When using subjects, reference them via @subject_name, e.g.: "@CharacterA and @CharacterB are having hotpot together"
audioBoolOptionalWhether to enable audio-video output. Default true for viduq3 and viduq3-turbo, false for other models.
audio_typeStringOptionalAudio type, effective when audio is true. Default all.
Options: all (sound effects + voice), speech_only (voice only), sound_effect_only (sound effects only)
durationIntOptionalVideo duration (seconds):
- viduq3-turbo, viduq3: Default 5, range 3–16
- viduq2-pro: Default 5, range 0–10 (0 for automatic duration)
- viduq2: Default 5, range 1–10
- viduq1: Default 5, fixed at 5
- vidu2.0: Default 4, fixed at 4
seedIntOptionalRandom seed. If omitted or set to 0, a random value is used.
aspect_ratioStringOptionalAspect ratio. Default 16:9. Options: 16:9, 9:16, 1:1.
Note: q2 models support arbitrary aspect ratios
resolutionStringOptionalResolution:
- viduq3-turbo, viduq3 (3–16s): Default 720p, options 540p, 720p, 1080p
- viduq2, viduq2-pro: Default 720p, options 540p, 720p, 1080p
- viduq1: Default 1080p, fixed at 1080p
- vidu2.0: Default 360p, options 360p, 720p
movement_amplitudeStringOptionalMovement amplitude. Default auto. Note: Ineffective for q2 and q3 models
off_peakBoolOptionalOff-peak mode. Default false.
Note: q3 models support off-peak when audio is true; q2/q1/2.0 series support off-peak when audio is false
watermarkBoolOptionalWhether to add a watermark. Default is no watermark.
wm_positionIntOptionalWatermark position. 1: Top-left, 2: Top-right, 3: Bottom-right (default), 4: Bottom-left
wm_urlStringOptionalCustom watermark image URL.
payloadStringOptionalPass-through parameter. Maximum 1048576 characters.
meta_dataStringOptionalMetadata identifier, JSON format string, pass-through field.
callback_urlStringOptionalCallback URL.

Request Example (Subject-Based)#


Mode 2: Non-Subject Invocation#

Directly pass reference images/videos, and the model automatically extracts subject consistency to generate the video.

Request Body#

ParameterTypeRequiredDescription
modelStringYesModel name. Options: viduq3-mix, viduq3-turbo, viduq3, viduq2-pro, viduq2, viduq1, vidu2.0.
- viduq3-mix: Strong visual quality, smart scene cutting, audio-video output, best overall balance
imagesArray[String]YesReference images. Supports 1–7 images (URL or Base64).
Note: viduq2-pro model supports max 1–4 images when uploading video
Supported formats: png, jpeg, jpg, webp; minimum resolution 128x128
videosArray[String]OptionalReference videos. Only supported by viduq2-pro.
Max 1 video of 8 seconds or 2 videos of 5 seconds.
Supported formats: mp4, avi, mov; max size 100M
promptStringYesText prompt. Maximum 2000 characters.
audioBoolOptionalWhether to enable audio-video output. Only q3 models support this in non-subject mode; default true.
bgmBoolOptionalWhether to add background music. Default false.
Note: Ineffective for q2 series when duration is 9s/10s; ineffective for q3 series
durationIntOptionalVideo duration (seconds):
- viduq3-turbo, viduq3-mix: Default 5, range 3–16
- viduq3: Default 5, range 3–16
- viduq2-pro: Default 5, range 0–10 (0 for automatic duration)
- viduq2: Default 5, range 1–10
- viduq1: Default 5, fixed at 5
- vidu2.0: Default 4, fixed at 4
seedIntOptionalRandom seed.
aspect_ratioStringOptionalAspect ratio. Default 16:9. Options: 16:9, 9:16, 4:3, 3:4, 1:1.
Note: 4:3 and 3:4 are only supported by q2 series models
resolutionStringOptionalResolution:
- viduq3-mix (3–16s): Default 720p, options 720p, 1080p
- viduq3-turbo (3–16s): Default 720p, options 540p, 720p, 1080p
- viduq3 (3–16s): Default 720p, options 540p, 720p, 1080p
- viduq2, viduq2-pro: Default 720p, options 540p, 720p, 1080p
- viduq1: Default 1080p, fixed at 1080p
- vidu2.0: Default 360p, options 360p, 720p
movement_amplitudeStringOptionalMovement amplitude. Default auto. Note: Ineffective for q2 and q3 series
off_peakBoolOptionalOff-peak mode. Default false.
Note: viduq3-mix does not support off-peak mode
watermarkBoolOptionalWhether to add a watermark.
wm_positionIntOptionalWatermark position.
wm_urlStringOptionalCustom watermark image URL.
payloadStringOptionalPass-through parameter.
meta_dataStringOptionalMetadata identifier.
callback_urlStringOptionalCallback URL.

Request Example (Non-Subject)#


Response Parameters#

FieldTypeDescription
task_idStringTask ID
stateStringProcessing state: created, queueing, processing, success, failed
modelStringModel name used for this call
promptStringPrompt
imagesArray[String]Image parameters
videosArray[String]Video parameters (returned for non-subject viduq2-pro calls)
durationIntVideo duration
seedIntRandom seed
aspect_ratioStringAspect ratio
resolutionStringResolution
bgmBoolWhether background music is added
audioBoolWhether audio-video output is enabled
audio_typeStringAudio type
movement_amplitudeStringMovement amplitude
payloadStringPass-through parameter
off_peakBoolWhether off-peak mode is used
creditsIntNumber of credits consumed by this call
watermarkBoolWhether a watermark is used
created_atStringTask creation time

Response Example#

{
  "task_id": "{task_id}",
  "state": "created",
  "model": "viduq3-mix",
  "images": ["https://example.com/ref1.png", "https://example.com/ref2.png"],
  "prompt": "Santa Claus and the bear hug by the lakeside.",
  "duration": 5,
  "seed": 123456,
  "aspect_ratio": "3:4",
  "resolution": "720p",
  "credits": 8,
  "created_at": "2025-01-01T15:41:31.968916Z"
}
Previous
02-Image-to-Video
Next
04-Start-End-Frame