the control pick. When the same character, product or voice has to survive across a series of clips, start here.
MiniMax H3 — the model everyone calls Hailuo 3.0 — landed on July 31st, and it changes the control equation: feed it up to 12 reference files (images, video clips, even voice recordings) and it generates a 2K clip with dialogue and sound baked in, keeping your character, your camera move, your voice. We wired it in the week it shipped.
Draft free on fast models · cheapest 2K on the site
The model is days old, so consider these first impressions from our own test renders rather than settled wisdom. We'll revise this section as the picture firms up — that's a feature of this page, not a disclaimer.
• Omni-reference control.
Up to 9 images, 3 video clips and 3 audio clips steer one generation. Character lock across a series of clips finally works without prompt gymnastics.
• 2K with sound, one pass.
Native 2560×1440 at 24 fps with dialogue, effects and ambience generated together. The output is publishable, not a silent draft.
• Instruction-based editing.
"Replace the car with a bicycle, keep everything else" edits the clip instead of re-rolling it. Early leaderboards rank H3 at the top for editing, and our tests agree.
• Price per pixel.
Early platform pricing puts a 15-second 2K clip at a fraction of what comparable models charge. The cheapest 2K we host.
• It's brand new.
Independent benchmark scores are still settling. Any "best model" claim this month — ours included — is provisional. Test before you standardize on it.
• Open weights: promised, not shipped.
MiniMax says weights are coming under a community license; as of early August none have. If self-hosting is your plan, wait for the download link, not the announcement.
• Reference curve.
Twelve input files is twelve ways to confuse it. Contradictory references (two different faces, clashing styles) produce muddy results — start with 2–3 and add.
• Naming chaos.
H3 = Hailuo 3.0 = Hailuo 03. It is not MiniMax M3 (a text model) and not Kling O3 (Kuaishou, unrelated). Worth knowing before you compare pricing pages.
H3 is a reference-first model — the files you attach matter as much as the words. Each recipe exercises a capability the other models here don't have.
Upload: 3 photos of the same person (front, profile, mid-shot)
"[Character] walks into a neon-lit ramen bar, sits at the counter, and orders. Handheld camera follows from behind, then cuts to a front shot across the counter. Warm interior light, street noise fading as the door closes."
Multiple angles of the same face give the model a 3D-ish understanding of the character. This is the setup for series work — reuse the same three photos across every clip and the character survives all of them.
Start from: a clip you already generated
"Replace the red sedan with a cargo bike. Keep the rider's clothing, the camera move, the rain, and all ambient sound exactly as they are."
H3 treats an existing clip as editable material rather than a failed attempt. Naming what must NOT change is the half of the instruction most people forget — it's what keeps the edit surgical.
Upload: 1 product photo + 1 short voice recording
"A presenter holds this exact product to camera in a bright studio and says, in the reference voice: 'Thirty days. No questions. Try it.' Cut to a close-up of the product on white. Upbeat room tone."
The audio reference locks the voice the way image references lock a face — a capability unique to H3 among our lineup. Two shots in one 15-second generation is also something the 5–8 second models simply can't do.
Same prompts across all three, judged by what we'd ship — with the caveat that H3 is new enough that these calls may age fast. The workflow stands regardless: draft cheap, compare, then commit credits, all from one prompt box.
Everyone in this category advertises "free unlimited." Ours has edges, and here they are:
previewed at WAIC on July 17, launched July 31. Native 2K, one-pass audio with dialogue, omni-reference, instruction editing, multi-shot clips. The version on this page.
the 1080p, ~10-second workhorse that built Hailuo's reputation for motion quality. No omni-reference, no native audio; what most "Hailuo" reviews you'll find were written about.
the generation that made the community take MiniMax seriously: objects fell, bounced and splashed convincingly at a price hobbyists could afford. The DNA is still visible in H3.
Same model, two names. MiniMax H3 is the official name; Hailuo 3.0 (sometimes written Hailuo 03) is what the community calls it, after the Hailuo app it runs in. It's the third generation of MiniMax's video line, launched July 31, 2026. Two things it is not: MiniMax M3, which is the company's text model, and Kling O3, a completely different model from Kuaishou that the 'Hailuo 03' spelling gets confused with.
The realist. When the clip has to pass as actual footage, with audio.
The choreographer. Best human motion and facial expression.
The sketchpad. Fastest, cheapest drafts — develop prompts here first.