MINIMAX H3 / MULTIMODAL INPUT
A precise brief.
A clear direction.
References. One request.
Combine image, sound and motion.
Connect.
Your key stays in memory only.
Build the context.
Images
0 / 9Character, product, style.
PNG / JPG / WEBP · UP TO 9Audio
0 / 3Voice, ambience, rhythm.
MP3 / WAV / M4A · UP TO 3Video
0 / 3Movement, framing, pacing.
MP4 / WEBM · UP TO 312 combined, not 15. Example: 6 images + 3 audio + 3 video.
At least one image or video. Clips: 2–15 seconds each; ≤15 seconds total per audio/video category.
Our upload limit: 25 MiB per file. Duration and decoding checks happen during preparation.
Select references above. Nothing uploads until you choose Upload selected.
Removing an uploaded reference excludes it from this request; it does not delete the stored file. Model input specifications ↗
Describe the scene.
This checks the request. Generation and upscaling are not enabled.
Track the request.
Waiting for a request
Connect with your client key to get started.
Status is the last recorded observation. No automatic polling or retries.
The details, unfiltered.
Ready when you are.
Connect to view an API response.