DataEyesAI
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
Official SiteConsoleDocs Home
Getting StartedDeveloper ToolsAI Models API
  1. MiniMax-H3 Video Generation
  • Getting Started
    • Overview
    • Console (Getting Started)
    • API Key
    • Base URL
  • Developer Tool Integration
    • OpenClaw
    • Claude Code
    • Codex
    • Gemini CLI
    • Grok CLI
    • Other Tools
  • AI Models API
    • OpenAI format (supports major original models)
      • Chat (Response)
        • Create Network Search
        • Create Model Response GPT-5 Enable Thinking
        • Create Function Call
        • Create Model Response
        • Create Model Response (Streaming Return)
        • Create Model Response (Control Thinking Length)
      • ChatGPT Interface
        • Audio
          • Audio to text gpt-4o-transcribe
          • GPT-4o-audio
          • Audio to text whisper-1
          • Audio to text gpt-4o-transcribe
          • Create voice gpt-4o-mini-tts
        • Chat
          • Create chat-based image recognition (non-streaming)
          • Create chat-based image recognition (streaming)
          • Create chat-based image recognition (streaming) best64
          • Official N test
          • Create structured output
          • Control the effort level of the inference model
          • Create chat function call
          • deepseek-ocr recognition
          • Create chat completion (non-stream)
        • Completions
          • ChatGPT automatic completion
          • Create completion
      • Image
        • Edit image
        • Create chat completion (streaming)
        • Create chat completion (qwen-mt-turbo)
        • Create chat completion with deepseek v3.1 level of reasoning (streaming)
      • Audio
        • Speech recognition
        • Speech synthesis
        • Official Function Calling invocation
        • Create chat-generated images (non-streaming)
      • Embedding
        • Text embeddings
    • Anthropic format
      • Chat
      • Chat(prompt cache)
      • Streaming response
      • Chat (deep reasoning)
      • Tool invocation (function call)
      • Analyze image
    • Google Gemini interface
      • Native format
        • Text-to-image + control over aspect ratio + clarity
        • Generate image
        • Text generation
        • Text generation - stream
        • Text generation + reasoning - stream
        • Image generation
        • Formatted output
        • Function call
        • Document understanding
        • URL context [native format]
        • Code execution
        • Video understanding
        • URL context
        • Video understanding - url [native format]
        • Imagen 4
        • Audio understanding
        • Embeddings
        • Chat
        • Edit image
      • Image-to-image Base64 request method
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity
        • Image editing
        • Single image gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image generation( gemini-2.5-flash-image)
        • Image generation gemini-2.5-flash-image, controlling aspect ratio.
        • Image understanding
      • Image-to-image URL request returns URL request format OpenAI
        • Single image generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Multi-image fusion slice generation with gemini-3-pro-image-preview, controlling aspect ratio and clarity.
        • Image understanding
    • NanoBanana
      • OpenAI request
        • Edit image
        • OpenAI image format
      • Gemini request
        • Generate image
        • Edit image
    • Midjourney format
      • Midjourney API Reference
      • Task query interface
      • Upload image
      • Get seed (Seed)
      • Submit Imagine task
      • Query tasks based on ID list
      • FaceSwap
      • Execute Action operation
      • /mj/submit/blend
      • Submit Describe task
      • Submit Modal
      • Refresh link
      • Edit image
      • Query task status by task ID
      • Get the seed of the task image
    • Doubao - Painting
      • doubao-seededit-3-0-i2i-250628
      • doubao-seedream-4-0-250828 - text-to-image
      • doubao-seedream-4-0-250828 - image-to-image
      • doubao-seedream-4-0-250828 - multi-image generation
    • Rerank Reordering Model
      • Rerank
    • Video Model
      • Grok Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-Video-Extension
      • Seedance Video Generation
        • 00-Overview
        • 01-Create-Video-Generation-Task
        • 02-Query-Video-Generation-Task
        • 03-Query-Video-Generation-Task-List
        • 04-Cancel-or-Delete-Task
        • Seedance Private Asset Library API Documentation
      • MiniMax-H3 Video Generation
        • 00-Overview
        • 01-Create-Video-Generation
        • 02-Create-Video-Regeneration
        • 03-Create-H3-Context-IR
        • 04-Query-Task
        • 05-List-Tasks
        • 06-Cancel-or-Delete-Task
      • Hailuo Video Generation
        • 00-Overview
        • 01-Text-to-Video-T2V
        • 02-Image-to-Video-I2V
        • 03-First-Last-Frame-FL2V
        • 04-Subject-Reference-S2V
        • 05-Query-Task-Status
        • 06-Video-Download
        • 99-Appendix-Camera-Movement-and-Webhooks
      • Jimeng Video Generation
        • 00-Overview
        • 01-3.0-Pro-Video-Generation
        • 02-720P-Text-to-Video
        • 03-720P-Image-to-Video-First-Frame
        • 04-720P-Image-to-Video-Start-End-Frame
        • 05-720P-Image-to-Video-Camera
        • 06-1080P-Text-to-Video
        • 07-1080P-Image-to-Video-First-Frame
        • 08-1080P-Image-to-Video-Start-End-Frame
        • 09-Error-Codes
      • Kling AI Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Omni-Video
        • 04-Multi-Image-to-Video
        • 05-Motion-Control
        • 06-Multi-Elements
        • 07-Video-Extension
        • 08-Lip-Sync
        • 09-Avatar
        • 10-Text-to-Audio
        • 11-Video-to-Audio
        • 12-TTS
        • 13-Custom-Voices
        • 14-Image-Recognition
        • 15-Element-Management
        • 16-Video-Effects
      • Vidu Video Generation
        • 00-Overview
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Start-End-Frame
        • 05-Multi-Frame
        • 06-Scene-Template
        • 07-Template-Story
        • 08-Query-Tasks
      • HappyHorse
        • HappyHorse Text-to-Video
        • HappyHorse Image-to-Video (First Frame)
        • HappyHorse Reference-to-Video
        • HappyHorse Video Editing
      • Wan Video Generation
        • 00-Overview.md
        • 01-Text-to-Video
        • 02-Image-to-Video
        • 03-Reference-to-Video
        • 04-Video-Editing
        • 05-First-Last-Frame-to-Video
        • 06-Motion-Transfer-and-Character-Swap
        • 07-Digital-Human-Video
        • 08-VACE-Video-Editing
        • 09-Query-Task
    • Audio API
      • Audio API
      • Gemini TTS API
      • Google DeepMind Lyria API
      • Elevenlabs Speech to Text API Reference
      • Text-to-Music Suno
        • Task Submission
          • Generate Song (Inspiration Mode)
          • Generate Song (Custom Mode)
          • Generate Song (Continuation Mode)
          • Generate Song (Singer Style)
          • Generate Song (Secondary Creation from Uploaded Song)
          • Generate Song (Song Stitching)
          • Generate Lyrics
          • Song Stitching
        • Query Interface
          • Batch Retrieve Tasks
          • Query Single Task
  • Search / Reader Product
    • Web Reader API​​
      • Web Reader API
      • Web Reader API(HK)
    • Web Search API​​
      • Modal Card API
        • Weather
          • All City ID
          • Weather Query API
      • Web Search API
      • Video Search api
      • Trending Search API
    • Document OCR Parsing API
      • fiel upload
      • URL Parsing
  • Advanced & System API
    • Data Updates
    • System interface
      • API Key & Quota Query API
      • API Key Management API
    • API Reference​​
      • Error Codes
      • HTTP Notes
    • List models
      • Models
  1. MiniMax-H3 Video Generation

01-Create-Video-Generation

Create Video Generation Task#

Version: v1.0.0 | Last Updated: 2026-08-13
This platform fully supports the MiniMax Video Generation V2 (MiniMax-H3) official API. Requests and responses are transparently proxied, and all parameter semantics remain consistent with the official API.
The model generates a video from multimodal inputs — text, images, video, and audio — with native 2K output. This is an asynchronous endpoint: on success it returns a task_id, which you poll via the Query Task endpoint.
POST https://platform.dataeyes.ai/hailuo/v2/video_generation

Request Parameters#

ParameterTypeRequiredDefaultDescription
modelstring (enum)Yes—Model name. Currently: MiniMax-H3
contentarrayYes—Multimodal input array (see below). Must include one non-empty text item
resolutionstring (enum)Yes—Video resolution: 768P or 2K
durationintegerYes—Output video duration in seconds, integer. Allowed values: 4-15
ratiostring (enum)ConditionaladaptiveAspect ratio; rules vary by scenario (see below)
callback_urlstringNo—Webhook URL for task status changes (see below)
aigc_watermarkbooleanNofalseWhether to add an AIGC watermark to the generated video

content Input Combinations#

Each element of the content array is typed via type (text / image_url / video_url / audio_url) and its purpose is marked via role. Different combinations map to different generation scenarios:
Scenariocontent combination
Text-to-video (t2va)A single text item
Image-to-video, first frametext + 1 image_url (role=first_frame; with a single image, role may be omitted and defaults to first frame)
Image-to-video, last frametext + 1 image_url (role=last_frame)
Image-to-video, first + last frametext + 2 image_url items (role set to first_frame and last_frame respectively)
Multimodal-reference-to-video (r2va)text + any combination of reference images (role=reference_image), reference videos (role=reference_video), and reference audio (role=reference_audio)
Image-to-video and multimodal reference are mutually exclusive: once any reference_image / reference_video / reference_audio role appears in content, first_frame / last_frame must not appear (and vice versa).

content Element Structure#

FieldTypeDescription
typestring (enum)text / image_url / video_url / audio_url
textstringPrompt text. Every request must include one non-empty text item. Up to 7,000 characters per text
image_url.urlstringImage address (required when type=image_url)
video_url.urlstringVideo address (required when type=video_url; multimodal reference only)
audio_url.urlstringAudio address (required when type=audio_url; multimodal reference only)
rolestring (enum)Purpose of the item, conditionally required: first_frame / last_frame / reference_image / reference_video / reference_audio
Each url accepts three forms:
A public URL;
mm_file://{file_id} (references an existing platform file, e.g. the file_id of a previous output);
A data:<mime>;base64,<Base64> data URI (lowercase mime subtype, e.g. data:image/png;base64,...).
The total request body must be ≤ 64 MB, and Base64 encoding inflates size by ~33%. Use public URLs or mm_file:// for large files instead of Base64.

Input Media Limits#

Images (image_url):
ItemLimit
FormatsJPG, JPEG, PNG, WEBP, HEIC, HEIF
Single file size≤ 30 MB
Width / height[256, 5760] px
Aspect ratio (w/h)[0.4, 2.5]
Countfirst frame ≤ 1, last frame ≤ 1, reference images ≤ 9
Videos (video_url, multimodal reference only):
ItemLimit
Container / formatMP4 (.mp4), MOV (.mov)
CodecsVideo: H.264/AVC, H.265/HEVC; audio: AAC, MP3
Single file size≤ 50 MB
Count≤ 3
Duration[2, 15] s per clip; total ≤ 15 s
Width / height[256, 5760] px
Aspect ratio (w/h)[0.4, 2.5]
Frame rate[23.976, 60]
Audio (audio_url, multimodal reference only):
ItemLimit
FormatsWAV, MP3
Single file size≤ 15 MB
Count≤ 3
Duration[2, 15] s per clip; total ≤ 15 s

ratio Rules#

Defaults to adaptive (the most suitable aspect ratio is chosen automatically from the inputs; the actual ratio is available in the query endpoint's ratio field). Allowed values: adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16.
ScenarioRule
Text-to-videoRequired, and must not be adaptive
Image-to-videoDetermined by the input image; always adaptive. Other values are ignored without error
Multimodal referenceOptional; defaults to adaptive, or specify any concrete ratio

callback_url Webhook#

When configured, the MiniMax server first sends a verification request containing a challenge field (echo the challenge back within 3 seconds to complete verification). After verification, every task status change is POSTed to the URL; the payload structure matches the Query Task response. Callback status values: queued / running / succeeded / failed / cancelled.

Request Examples#

Text-to-video:
Image-to-video (first frame):
Multimodal-reference-to-video:

Response Example#

{
  "task_id": "424010985738629"
}
FieldTypeDescription
task_idstringTask ID; use it with Query Task to retrieve status and results

Billing#

Billed per second on usage.total_seconds (input reference video seconds + output seconds) at the resolution's unit price; the first 5 input images per task are free, additional images are billed per image. An estimated fee is pre-charged at submission and settled against actual usage on completion (difference charged or refunded); failed tasks are fully refunded. See Overview · Billing & Settlement.
Previous
00-Overview
Next
02-Create-Video-Regeneration