DataEyesAI
Official SiteConsoleDocs HomeGetting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs HomeGetting StartedDeveloper ToolsAI Models API
  1. MiniMax-H3 Video Generation
  • OpenAI format (supports major original models)
    • Chat (Response)
      • Create Network Search
      • Create Model Response GPT-5 Enable Thinking
      • Create Function Call
      • Create Model Response
      • Create Model Response (Streaming Return)
      • Create Model Response (Control Thinking Length)
    • ChatGPT Interface
      • Audio
        • Audio to text gpt-4o-transcribe
        • GPT-4o-audio
        • Audio to text whisper-1
        • Audio to text gpt-4o-transcribe
        • Create voice gpt-4o-mini-tts
      • Chat
        • Create chat-based image recognition (non-streaming)
        • Create chat-based image recognition (streaming)
        • Create chat-based image recognition (streaming) best64
        • Official N test
        • Create structured output
        • Control the effort level of the inference model
        • Create chat function call
        • deepseek-ocr recognition
        • Create chat completion (non-stream)
      • Completions
        • ChatGPT automatic completion
        • Create completion
    • Image
      • Edit image
      • Create chat completion (streaming)
      • Create chat completion (qwen-mt-turbo)
      • Create chat completion with deepseek v3.1 level of reasoning (streaming)
    • Audio
      • Speech recognition
      • Speech synthesis
      • Official Function Calling invocation
      • Create chat-generated images (non-streaming)
    • Embedding
      • Text embeddings
  • Anthropic format
    • Chat
    • Chat(prompt cache)
    • Streaming response
    • Chat (deep reasoning)
    • Tool invocation (function call)
    • Analyze image
  • Google Gemini interface
    • Native format
      • Text-to-image + control over aspect ratio + clarity
      • Generate image
      • Text generation
      • Text generation - stream
      • Text generation + reasoning - stream
      • Image generation
      • Formatted output
      • Function call
      • Document understanding
      • URL context [native format]
      • Code execution
      • Video understanding
      • URL context
      • Video understanding - url [native format]
      • Imagen 4
      • Audio understanding
      • Embeddings
      • Chat
      • Edit image
    • Image-to-image Base64 request method
      • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
      • Image editing
      • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Image generation( gemini-2.5-flash-image)
      • Image generation gemini-2.5-flash-image, controlling aspect ratio.
      • Image understanding
    • Image-to-image URL request returns URL request format OpenAI
      • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
      • Image understanding
  • NanoBanana
    • OpenAI request
      • Edit image
      • OpenAI image format
    • Gemini request
      • Generate image
      • Edit image
  • Midjourney format
    • Midjourney API Reference
    • Task query interface
    • Upload image
    • Get seed (Seed)
    • Submit Imagine task
    • Query tasks based on ID list
    • FaceSwap
    • Execute Action operation
    • /mj/submit/blend
    • Submit Describe task
    • Submit Modal
    • Refresh link
    • Edit image
    • Query task status by task ID
    • Get the seed of the task image
  • Doubao - Painting
    • doubao-seededit-3-0-i2i-250628
    • doubao-seedream-4-0-250828 - text-to-image
    • doubao-seedream-4-0-250828 - image-to-image
    • doubao-seedream-4-0-250828 - multi-image generation
  • Rerank Reordering Model
    • Rerank
  • Video Model
    • Grok Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Video-Editing
      • 05-Video-Extension
    • Seedance Video Generation
      • 00-Overview
      • 01-Create-Video-Generation-Task
      • 02-Query-Video-Generation-Task
      • 03-Query-Video-Generation-Task-List
      • 04-Cancel-or-Delete-Task
      • Seedance Private Asset Library API Documentation
    • MiniMax-H3 Video Generation
      • 00-Overview
      • 01-Create-Video-Generation
      • 02-Create-Video-Regeneration
      • 03-Create-H3-Context-IR
      • 04-Query-Task
      • 05-List-Tasks
      • 06-Cancel-or-Delete-Task
    • Hailuo Video Generation
      • 00-Overview
      • 01-Text-to-Video-T2V
      • 02-Image-to-Video-I2V
      • 03-First-Last-Frame-FL2V
      • 04-Subject-Reference-S2V
      • 05-Query-Task-Status
      • 06-Video-Download
      • 99-Appendix-Camera-Movement-and-Webhooks
    • Jimeng Video Generation
      • 00-Overview
      • 01-3.0-Pro-Video-Generation
      • 02-720P-Text-to-Video
      • 03-720P-Image-to-Video-First-Frame
      • 04-720P-Image-to-Video-Start-End-Frame
      • 05-720P-Image-to-Video-Camera
      • 06-1080P-Text-to-Video
      • 07-1080P-Image-to-Video-First-Frame
      • 08-1080P-Image-to-Video-Start-End-Frame
      • 09-Error-Codes
    • Kling AI Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Omni-Video
      • 04-Multi-Image-to-Video
      • 05-Motion-Control
      • 06-Multi-Elements
      • 07-Video-Extension
      • 08-Lip-Sync
      • 09-Avatar
      • 10-Text-to-Audio
      • 11-Video-to-Audio
      • 12-TTS
      • 13-Custom-Voices
      • 14-Image-Recognition
      • 15-Element-Management
      • 16-Video-Effects
    • Vidu Video Generation
      • 00-Overview
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Start-End-Frame
      • 05-Multi-Frame
      • 06-Scene-Template
      • 07-Template-Story
      • 08-Query-Tasks
    • HappyHorse
      • HappyHorse Text-to-Video
      • HappyHorse Image-to-Video (First Frame)
      • HappyHorse Reference-to-Video
      • HappyHorse Video Editing
    • Wan Video Generation
      • 00-Overview.md
      • 01-Text-to-Video
      • 02-Image-to-Video
      • 03-Reference-to-Video
      • 04-Video-Editing
      • 05-First-Last-Frame-to-Video
      • 06-Motion-Transfer-and-Character-Swap
      • 07-Digital-Human-Video
      • 08-VACE-Video-Editing
      • 09-Query-Task
  • Audio API
    • Audio API
    • Gemini TTS API
    • Google DeepMind Lyria API
    • Elevenlabs Speech to Text API Reference
    • Text-to-Music Suno
      • Task Submission
        • Generate Song (Inspiration Mode)
        • Generate Song (Custom Mode)
        • Generate Song (Continuation Mode)
        • Generate Song (Singer Style)
        • Generate Song (Secondary Creation from Uploaded Song)
        • Generate Song (Song Stitching)
        • Generate Lyrics
        • Song Stitching
      • Query Interface
        • Batch Retrieve Tasks
        • Query Single Task
  1. MiniMax-H3 Video Generation

01-Create-Video-Generation

Create Video Generation Task#

Version: v1.0.0 | Last Updated: 2026-08-13
This platform fully supports the MiniMax Video Generation V2 (MiniMax-H3) official API. Requests and responses are transparently proxied, and all parameter semantics remain consistent with the official API.
The model generates a video from multimodal inputs — text, images, video, and audio — with native 2K output. This is an asynchronous endpoint: on success it returns a task_id, which you poll via the Query Task endpoint.
POST https://platform.dataeyes.ai/hailuo/v2/video_generation

Request Parameters#

ParameterTypeRequiredDefaultDescription
modelstring (enum)Yes—Model name. Currently: MiniMax-H3
contentarrayYes—Multimodal input array (see below). Must include one non-empty text item
resolutionstring (enum)Yes—Video resolution: 768P or 2K
durationintegerYes—Output video duration in seconds, integer. Allowed values: 4-15
ratiostring (enum)ConditionaladaptiveAspect ratio; rules vary by scenario (see below)
callback_urlstringNo—Webhook URL for task status changes (see below)
aigc_watermarkbooleanNofalseWhether to add an AIGC watermark to the generated video

content Input Combinations#

Each element of the content array is typed via type (text / image_url / video_url / audio_url) and its purpose is marked via role. Different combinations map to different generation scenarios:
Scenariocontent combination
Text-to-video (t2va)A single text item
Image-to-video, first frametext + 1 image_url (role=first_frame; with a single image, role may be omitted and defaults to first frame)
Image-to-video, last frametext + 1 image_url (role=last_frame)
Image-to-video, first + last frametext + 2 image_url items (role set to first_frame and last_frame respectively)
Multimodal-reference-to-video (r2va)text + any combination of reference images (role=reference_image), reference videos (role=reference_video), and reference audio (role=reference_audio)
Image-to-video and multimodal reference are mutually exclusive: once any reference_image / reference_video / reference_audio role appears in content, first_frame / last_frame must not appear (and vice versa).

content Element Structure#

FieldTypeDescription
typestring (enum)text / image_url / video_url / audio_url
textstringPrompt text. Every request must include one non-empty text item. Up to 7,000 characters per text
image_url.urlstringImage address (required when type=image_url)
video_url.urlstringVideo address (required when type=video_url; multimodal reference only)
audio_url.urlstringAudio address (required when type=audio_url; multimodal reference only)
rolestring (enum)Purpose of the item, conditionally required: first_frame / last_frame / reference_image / reference_video / reference_audio
Each url accepts three forms:
A public URL;
mm_file://{file_id} (references an existing platform file, e.g. the file_id of a previous output);
A data:<mime>;base64,<Base64> data URI (lowercase mime subtype, e.g. data:image/png;base64,...).
The total request body must be ≤ 64 MB, and Base64 encoding inflates size by ~33%. Use public URLs or mm_file:// for large files instead of Base64.

Input Media Limits#

Images (image_url):
ItemLimit
FormatsJPG, JPEG, PNG, WEBP, HEIC, HEIF
Single file size≤ 30 MB
Width / height[256, 5760] px
Aspect ratio (w/h)[0.4, 2.5]
Countfirst frame ≤ 1, last frame ≤ 1, reference images ≤ 9
Videos (video_url, multimodal reference only):
ItemLimit
Container / formatMP4 (.mp4), MOV (.mov)
CodecsVideo: H.264/AVC, H.265/HEVC; audio: AAC, MP3
Single file size≤ 50 MB
Count≤ 3
Duration[2, 15] s per clip; total ≤ 15 s
Width / height[256, 5760] px
Aspect ratio (w/h)[0.4, 2.5]
Frame rate[23.976, 60]
Audio (audio_url, multimodal reference only):
ItemLimit
FormatsWAV, MP3
Single file size≤ 15 MB
Count≤ 3
Duration[2, 15] s per clip; total ≤ 15 s

ratio Rules#

Defaults to adaptive (the most suitable aspect ratio is chosen automatically from the inputs; the actual ratio is available in the query endpoint's ratio field). Allowed values: adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16.
ScenarioRule
Text-to-videoRequired, and must not be adaptive
Image-to-videoDetermined by the input image; always adaptive. Other values are ignored without error
Multimodal referenceOptional; defaults to adaptive, or specify any concrete ratio

callback_url Webhook#

When configured, the MiniMax server first sends a verification request containing a challenge field (echo the challenge back within 3 seconds to complete verification). After verification, every task status change is POSTed to the URL; the payload structure matches the Query Task response. Callback status values: queued / running / succeeded / failed / cancelled.

Request Examples#

Text-to-video:
Image-to-video (first frame):
Multimodal-reference-to-video:

Response Example#

{
  "task_id": "424010985738629"
}
FieldTypeDescription
task_idstringTask ID; use it with Query Task to retrieve status and results

Billing#

Billed per second on usage.total_seconds (input reference video seconds + output seconds) at the resolution's unit price; the first 5 input images per task are free, additional images are billed per image. An estimated fee is pre-charged at submission and settled against actual usage on completion (difference charged or refunded); failed tasks are fully refunded. See Overview · Billing & Settlement.
Previous
00-Overview
Next
02-Create-Video-Regeneration