A new-generation video model from MiniMax: picture and audio are generated at the same time, so no separate voice-over pass is needed. The prompt describes not only the scene and camera move but the soundtrack too — music, ambience, effects. Text-to-video and image-to-video, 5–15 seconds, 480p and 768p, 7 aspect ratios.
Cinematic aerial push-in over a stormy ocean at dusk, huge waves crashing against dark cliffs, lightning flashing on the horizon, rolling thunder and a deep orchestral score
In the independent Artificial Analysis ranking MiniMax H3 sits first in the video editing category (Elo 1132) and second in text-to-video (Elo 1238), behind only Gemini Omni Flash. Among open-weight models H3 leads both categories. The editing lead is slim — nine Elo points against a confidence interval of about ten, so it is effectively a tie with the runner-up, and the ranking is still young: H3 has fewer votes than its neighbours in the table.
Artificial Analysis Video Arena, measured 06.08.2026 · model released 31.07.2026, weights published 03.08.2026
What the model takes in and what it returns
768p is the model's native canvas; 480p is faster and costs half as much. A 16:9 frame at 768p comes out at 1344×768 pixels.
Duration is set in whole seconds. A five-second 480p clip is usually ready in under a minute and a half.
16:9, 9:16, 1:1, 4:3, 3:4, 21:9 and 9:21. In image-to-video the ratio comes from the source frame.
Audio arrives inside the finished MP4 as a separate two-channel AAC track. There is no toggle for it — sound is always generated.
Text-to-video and image-to-video at 24 fps. The second mode also accepts a final frame for a controlled transition.
The output is a ready MP4 carrying both the video and the audio track — no extra assembly or mixing step.
Four scenarios the model is actually picked for
The stereo track is generated in the same pass as the picture. Describe the soundtrack you want right in the prompt — "rolling thunder and a deep orchestral score", "quiet jazz and a soft cafe hum" — and it lands in the finished clip in sync with the motion.
A street saxophonist plays under a subway underpass, warm light from above, commuters passing in soft blur, the saxophone melody echoing off tiled walls with a distant train rumbling past
Scene, camera move and sound from a single description, with no source material. Seven aspect ratios are available, including vertical 9:16 for social and wide 21:9 for cinematic shots.
Cinematic drone shot descending over a Marrakesh spice market at golden hour, vivid mounds of saffron and paprika, vendors calling out, sizzling street food and rhythmic hand percussion
Your image becomes the first frame and the aspect ratio is inherited from it, so nothing needs reformatting. The prompt drives the motion inside the scene and the sound.
The fisherman blinks and turns his head slightly toward the sea, wind moves his hair, seagulls cry overhead and waves break against the wooden pier behind him
You can supply a second image as the final frame and the model builds the transition between them. Handy for scene changes, transformations and loopable clips.
Time passes on the same cafe terrace: dawn mist fades, lamps and window lights switch on one by one, chairs unfold around the tables and guests settle in, the quiet morning street gradually fills with evening chatter and soft accordion music
Three steps to your first clip with sound
Text-to-video if you start from scratch. Image-to-video if you already have a frame you want to bring to life.
Beyond the action and camera move, add the soundtrack you want: music genre, ambience, effects. Prompts run up to 7000 characters and eleven languages are officially supported.
Choose resolution and duration — the cost in credits updates immediately. The finished clip appears in your account with the audio already inside.
Tasks where built-in audio saves a whole production step
Vertical 9:16 and a finished audio track — the clip does not need a separate sound pass before publishing.
Product shots with atmospheric sound: a liquid splash, city noise, a music bed matching the mood you need.
A still frame turns into a short scene with motion and ambient sound, keeping the original aspect ratio.
Nature, weather, cityscapes — shots where half the impression comes from the sound itself: rain, thunder, wind, surf.
Run the model without setting up an environment
Nothing to install or configure: open the page, pick a mode and start generating.
The cost is visible before you start and recalculates when you change resolution or duration.
Finished clips can be used in commercial projects — ads, social media and client work.
Short answers to what people ask most about the model
MiniMax H3 is a video model from MiniMax released on July 31, 2026. What sets it apart from the previous generation is that picture and sound are generated simultaneously, in a single pass: the finished MP4 already carries a stereo track. The model works in two modes — from a text description and from an uploaded image — and produces clips of 5 to 15 seconds at 480p or 768p.
The model was built by MiniMax. Its official name is MiniMax H3, without the word Hailuo: the Hailuo line stayed at version 2.3, while H3 is the third generation and a separate family. The model shipped on July 31, 2026, and its weights were published openly on August 3 — still uncommon for a video model at this level.
The audio track is produced in the same pass as the video and arrives inside the finished MP4 as 32 kHz stereo. There is no "add sound" button — audio is always generated. The developer stresses that speech, effects and music are modelled jointly, without splitting them into separate domains, which is why the sound matches the motion in frame: thunder lands on the lightning flash, a splash is heard the moment the liquid hits.
Two. Text-to-video builds the scene from scratch out of a description and lets you pick the aspect ratio. Image-to-video takes your frame as the first one and animates it, inheriting the aspect ratio from the source. The second mode also accepts a final frame, giving you a transition between two images.
480p and 768p. 768p is the model's native canvas — at 16:9 that is a 1344×768 frame. The 480p mode is noticeably faster and costs half as much, which makes it convenient for draft runs and prompt tuning, leaving 768p for the final take.
From 5 to 15 seconds, set in whole seconds. Generation time depends on resolution: a five-second 480p clip is usually ready in under a minute and a half, while 768p takes longer.
Seven: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 and 9:21. The vertical ones suit social platforms, wide 21:9 suits cinematic shots. In image-to-video you do not pick a ratio — it is inherited from the uploaded frame.
Yes. In image-to-video you can upload a second image alongside the first, and the model builds the transition between them. This is useful for transformations, scene changes and looping clips where the start and the end need to match.
Sound goes into the same text as the picture — there is no separate field for it. Phrasing like "rolling thunder and a deep orchestral score", "quiet jazz and a muted cafe hum" or "surf and seagull calls" works well. If you do not describe the sound, the model picks it itself based on the scene. Prompts run up to 7000 characters.
A single clip caps at 15 seconds and 768p — 4K requires a separate upscale pass. Like other video models, H3 struggles to hold small text in frame: lettering on objects usually distorts when an image is animated, so text is better overlaid after generation. The developer also notes that visual detail and multimodal understanding still have room to improve.