DataEyesAI
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models APITerms & Policies
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models APITerms & Policies
  1. MiniMax-H3 Video Generation
  • Documentation
    • Getting Started
      • Overview
      • Console (Getting Started)
      • API Key
      • Base URL
    • Developer Tool Integration
      • OpenClaw
      • Claude Code
      • Codex
      • Gemini CLI
      • Grok CLI
      • Other Tools
    • AI Models API
      • OpenAI format (supports major original models)
        • Chat (Response)
          • Create Network Search
          • Create Model Response GPT-5 Enable Thinking
          • Create Function Call
          • Create Model Response
          • Create Model Response (Streaming Return)
          • Create Model Response (Control Thinking Length)
        • ChatGPT Interface
          • Audio
            • Audio to text gpt-4o-transcribe
            • GPT-4o-audio
            • Audio to text whisper-1
            • Audio to text gpt-4o-transcribe
            • Create voice gpt-4o-mini-tts
          • Chat
            • Create chat-based image recognition (non-streaming)
            • Create chat-based image recognition (streaming)
            • Create chat-based image recognition (streaming) best64
            • Official N test
            • Create structured output
            • Control the effort level of the inference model
            • Create chat function call
            • deepseek-ocr recognition
            • Create chat completion (non-stream)
          • Completions
            • ChatGPT automatic completion
            • Create completion
        • Image
          • Edit image
          • Create chat completion (streaming)
          • Create chat completion (qwen-mt-turbo)
          • Create chat completion with deepseek v3.1 level of reasoning (streaming)
        • Audio
          • Speech recognition
          • Speech synthesis
          • Official Function Calling invocation
          • Create chat-generated images (non-streaming)
        • Embedding
          • Text embeddings
      • Anthropic format
        • Chat
        • Chat(prompt cache)
        • Streaming response
        • Chat (deep reasoning)
        • Tool invocation (function call)
        • Analyze image
      • Google Gemini interface
        • Native format
          • Text-to-image + control over aspect ratio + clarity
          • Generate image
          • Text generation
          • Text generation - stream
          • Text generation + reasoning - stream
          • Image generation
          • Formatted output
          • Function call
          • Document understanding
          • URL context [native format]
          • Code execution
          • Video understanding
          • URL context
          • Video understanding - url [native format]
          • Imagen 4
          • Audio understanding
          • Embeddings
          • Chat
          • Edit image
        • Image-to-image Base64 request method
          • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
          • Image editing
          • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
          • Image generation( gemini-2.5-flash-image)
          • Image generation gemini-2.5-flash-image, controlling aspect ratio.
          • Image understanding
        • Image-to-image URL request returns URL request format OpenAI
          • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
          • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
          • Image understanding
      • NanoBanana
        • OpenAI request
          • Edit image
          • OpenAI image format
        • Gemini request
          • Generate image
          • Edit image
      • Midjourney format
        • Midjourney API Reference
        • Task query interface
        • Upload image
        • Get seed (Seed)
        • Submit Imagine task
        • Query tasks based on ID list
        • FaceSwap
        • Execute Action operation
        • /mj/submit/blend
        • Submit Describe task
        • Submit Modal
        • Refresh link
        • Edit image
        • Query task status by task ID
        • Get the seed of the task image
      • Doubao - Painting
        • doubao-seededit-3-0-i2i-250628
        • doubao-seedream-4-0-250828 - text-to-image
        • doubao-seedream-4-0-250828 - image-to-image
        • doubao-seedream-4-0-250828 - multi-image generation
      • Rerank Reordering Model
        • Rerank
      • Video Model
        • Grok Video Generation
          • 00-Overview
          • 01-Text-to-Video
          • 02-Image-to-Video
          • 03-Reference-to-Video
          • 04-Video-Editing
          • 05-Video-Extension
        • Seedance Video Generation
          • 00-Overview
          • 01-Create-Video-Generation-Task
          • 02-Query-Video-Generation-Task
          • 03-Query-Video-Generation-Task-List
          • 04-Cancel-or-Delete-Task
          • Seedance Private Asset Library API Documentation
        • MiniMax-H3 Video Generation
          • 00-Overview
          • 01-Create-Video-Generation
          • 02-Create-Video-Regeneration
          • 03-Create-H3-Context-IR
          • 04-Query-Task
          • 05-List-Tasks
          • 06-Cancel-or-Delete-Task
        • Hailuo Video Generation
          • 00-Overview
          • 01-Text-to-Video-T2V
          • 02-Image-to-Video-I2V
          • 03-First-Last-Frame-FL2V
          • 04-Subject-Reference-S2V
          • 05-Query-Task-Status
          • 06-Video-Download
          • 99-Appendix-Camera-Movement-and-Webhooks
        • Jimeng Video Generation
          • 00-Overview
          • 01-3.0-Pro-Video-Generation
          • 02-720P-Text-to-Video
          • 03-720P-Image-to-Video-First-Frame
          • 04-720P-Image-to-Video-Start-End-Frame
          • 05-720P-Image-to-Video-Camera
          • 06-1080P-Text-to-Video
          • 07-1080P-Image-to-Video-First-Frame
          • 08-1080P-Image-to-Video-Start-End-Frame
          • 09-Error-Codes
        • Kling AI Video Generation
          • 00-Overview
          • 01-Text-to-Video
          • 02-Image-to-Video
          • 03-Omni-Video
          • 04-Multi-Image-to-Video
          • 05-Motion-Control
          • 06-Multi-Elements
          • 07-Video-Extension
          • 08-Lip-Sync
          • 09-Avatar
          • 10-Text-to-Audio
          • 11-Video-to-Audio
          • 12-TTS
          • 13-Custom-Voices
          • 14-Image-Recognition
          • 15-Element-Management
          • 16-Video-Effects
        • Vidu Video Generation
          • 00-Overview
          • 01-Text-to-Video
          • 02-Image-to-Video
          • 03-Reference-to-Video
          • 04-Start-End-Frame
          • 05-Multi-Frame
          • 06-Scene-Template
          • 07-Template-Story
          • 08-Query-Tasks
        • HappyHorse
          • HappyHorse Text-to-Video
          • HappyHorse Image-to-Video (First Frame)
          • HappyHorse Reference-to-Video
          • HappyHorse Video Editing
        • Wan Video Generation
          • 00-Overview.md
          • 01-Text-to-Video
          • 02-Image-to-Video
          • 03-Reference-to-Video
          • 04-Video-Editing
          • 05-First-Last-Frame-to-Video
          • 06-Motion-Transfer-and-Character-Swap
          • 07-Digital-Human-Video
          • 08-VACE-Video-Editing
          • 09-Query-Task
      • Audio API
        • Audio API
        • Gemini TTS API
        • Google DeepMind Lyria API
        • Elevenlabs Speech to Text API Reference
        • Text-to-Music Suno
          • Task Submission
            • Generate Song (Inspiration Mode)
            • Generate Song (Custom Mode)
            • Generate Song (Continuation Mode)
            • Generate Song (Singer Style)
            • Generate Song (Secondary Creation from Uploaded Song)
            • Generate Song (Song Stitching)
            • Generate Lyrics
            • Song Stitching
          • Query Interface
            • Batch Retrieve Tasks
            • Query Single Task
    • Search / Reader Product
      • Web Reader API​​
        • Web Reader API
        • Web Reader API(HK)
      • Web Search API​​
        • Modal Card API
          • Weather
            • All City ID
            • Weather Query API
        • Web Search API
        • Video Search api
        • Trending Search API
      • Document OCR Parsing API
        • fiel upload
        • URL Parsing
    • Advanced & System API
      • Data Updates
      • System interface
        • API Key & Quota Query API
        • API Key Management API
      • API Reference​​
        • Error Codes
        • HTTP Notes
      • List models
        • Models
  • Terms & Policies
    • DataEyesAI API Terms of Service
    • DataEyesAI Legal Notice and Privacy Policy
    • DataEyesAI Paid Services Agreement
    • Automatic Renewal Service Rules
  1. MiniMax-H3 Video Generation

01-Create-Video-Generation

Create Video Generation Task#

Version: v1.0.0 | Last Updated: 2026-08-13
This platform fully supports the MiniMax Video Generation V2 (MiniMax-H3) official API. Requests and responses are transparently proxied, and all parameter semantics remain consistent with the official API.
The model generates a video from multimodal inputs — text, images, video, and audio — with native 2K output. This is an asynchronous endpoint: on success it returns a task_id, which you poll via the Query Task endpoint.
POST https://platform.dataeyes.ai/hailuo/v2/video_generation

Request Parameters#

ParameterTypeRequiredDefaultDescription
modelstring (enum)Yes—Model name. Currently: MiniMax-H3
contentarrayYes—Multimodal input array (see below). Must include one non-empty text item
resolutionstring (enum)Yes—Video resolution: 768P or 2K
durationintegerYes—Output video duration in seconds, integer. Allowed values: 4-15
ratiostring (enum)ConditionaladaptiveAspect ratio; rules vary by scenario (see below)
callback_urlstringNo—Webhook URL for task status changes (see below)
aigc_watermarkbooleanNofalseWhether to add an AIGC watermark to the generated video

content Input Combinations#

Each element of the content array is typed via type (text / image_url / video_url / audio_url) and its purpose is marked via role. Different combinations map to different generation scenarios:
Scenariocontent combination
Text-to-video (t2va)A single text item
Image-to-video, first frametext + 1 image_url (role=first_frame; with a single image, role may be omitted and defaults to first frame)
Image-to-video, last frametext + 1 image_url (role=last_frame)
Image-to-video, first + last frametext + 2 image_url items (role set to first_frame and last_frame respectively)
Multimodal-reference-to-video (r2va)text + any combination of reference images (role=reference_image), reference videos (role=reference_video), and reference audio (role=reference_audio)
Image-to-video and multimodal reference are mutually exclusive: once any reference_image / reference_video / reference_audio role appears in content, first_frame / last_frame must not appear (and vice versa).

content Element Structure#

FieldTypeDescription
typestring (enum)text / image_url / video_url / audio_url
textstringPrompt text. Every request must include one non-empty text item. Up to 7,000 characters per text
image_url.urlstringImage address (required when type=image_url)
video_url.urlstringVideo address (required when type=video_url; multimodal reference only)
audio_url.urlstringAudio address (required when type=audio_url; multimodal reference only)
rolestring (enum)Purpose of the item, conditionally required: first_frame / last_frame / reference_image / reference_video / reference_audio
Each url accepts three forms:
A public URL;
mm_file://{file_id} (references an existing platform file, e.g. the file_id of a previous output);
A data:<mime>;base64,<Base64> data URI (lowercase mime subtype, e.g. data:image/png;base64,...).
The total request body must be ≤ 64 MB, and Base64 encoding inflates size by ~33%. Use public URLs or mm_file:// for large files instead of Base64.

Input Media Limits#

Images (image_url):
ItemLimit
FormatsJPG, JPEG, PNG, WEBP, HEIC, HEIF
Single file size≤ 30 MB
Width / height[256, 5760] px
Aspect ratio (w/h)[0.4, 2.5]
Countfirst frame ≤ 1, last frame ≤ 1, reference images ≤ 9
Videos (video_url, multimodal reference only):
ItemLimit
Container / formatMP4 (.mp4), MOV (.mov)
CodecsVideo: H.264/AVC, H.265/HEVC; audio: AAC, MP3
Single file size≤ 50 MB
Count≤ 3
Duration[2, 15] s per clip; total ≤ 15 s
Width / height[256, 5760] px
Aspect ratio (w/h)[0.4, 2.5]
Frame rate[23.976, 60]
Audio (audio_url, multimodal reference only):
ItemLimit
FormatsWAV, MP3
Single file size≤ 15 MB
Count≤ 3
Duration[2, 15] s per clip; total ≤ 15 s

ratio Rules#

Defaults to adaptive (the most suitable aspect ratio is chosen automatically from the inputs; the actual ratio is available in the query endpoint's ratio field). Allowed values: adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16.
ScenarioRule
Text-to-videoRequired, and must not be adaptive
Image-to-videoDetermined by the input image; always adaptive. Other values are ignored without error
Multimodal referenceOptional; defaults to adaptive, or specify any concrete ratio

callback_url Webhook#

When configured, the MiniMax server first sends a verification request containing a challenge field (echo the challenge back within 3 seconds to complete verification). After verification, every task status change is POSTed to the URL; the payload structure matches the Query Task response. Callback status values: queued / running / succeeded / failed / cancelled.

Request Examples#

Text-to-video:
Image-to-video (first frame):
Multimodal-reference-to-video:

Response Example#

{
  "task_id": "424010985738629"
}
FieldTypeDescription
task_idstringTask ID; use it with Query Task to retrieve status and results

Billing#

Billed per second on usage.total_seconds (input reference video seconds + output seconds) at the resolution's unit price; the first 5 input images per task are free, additional images are billed per image. An estimated fee is pre-charged at submission and settled against actual usage on completion (difference charged or refunded); failed tasks are fully refunded. See Overview · Billing & Settlement.
Previous
00-Overview
Next
02-Create-Video-Regeneration