Kuaishou's reference-first model: your images and clips define the subject, MVL keeps it identical shot after shot — and you edit the result by talking to it. Type a prompt below and try it.
Developer
Kuaishou
Type
Reference-led multimodal (MVL)
References
Up to 7 images + video
Duration
3–15 s per clip
Audio
Native · dialogue in 5 languages
Editing
Conversational, in-place
Tier here
Paid credits (Std / Pro)
Our one-line verdict: the consistency pick for series work. When the same subject must survive dozens of generations — and you'd rather fix clips than re-roll them — start here.
Kling O3 — Video 3.0 Omni, the reference-led half of Kuaishou's 3.0 generation — treats your uploads as ground truth. Feed it up to seven images of a subject, optionally a video clip, and its MVL architecture keeps that exact face, logo or prop stable through camera moves, scene changes and multi-shot sequences; independent reviewers called the O3 series the most controllable AI video of early 2026. The other half of its case: conversational editing — 'remove the passerby, keep everything else' — applied to finished clips instead of re-rolling them. One spelling note that sends people to the wrong page: Kling O3 is not Hailuo 03. That's MiniMax H3, a different company's model.
Judged from our own renders and the published record, revised as the facts move — that's a feature of this page, not a disclaimer.
• Reference-locked identity.
Up to 7 images plus optional video define the subject; MVL keeps faces, clothing details, logos and text stable through complex camera movement. The series-work benchmark.
• Conversational video editing.
Object removal and replacement, background changes, effects — issued as natural-language commands against a finished clip. Fixing beats re-rolling.
• Multi-character coreference.
Several referenced characters in one scene, each holding their own identity — ensemble shots that single-reference models scramble.
• Voice from your references.
Custom voices derived from video or image inputs, with native dialogue in five languages — a brand spokesperson who sounds the same in every clip.
• Reference count mid-pack.
7 images + video is plenty for one subject, but MiniMax H3 takes 12 files and Seedance 2.5 takes 50 — full asset-kit productions may want the bigger intake.
• The naming trap.
'Kling O3' and 'Hailuo 03' are different companies' models; O3 vs V3 vs 2.6 confuses even careful buyers. Check the model chip before you spend.
• 1080p standard.
The O3 line's standard output sits at HD; the family's 4K lives on other variants. Resolution briefs go to Kling 3.0 or MiniMax H3.
• Prompt-only work is the twin's job.
With no references to anchor, its edge narrows — pure text-to-video briefs usually come out better on prompt-first Kling 3.0.
Copy into the generator above, swap the brackets, render.
Upload: 7 photos of the same person (angles, expressions, full-body)
"The character walks through a morning market, then appears at an office desk, then on a rooftop at sunset — same face, same build, outfits changing per scene. One spoken line in each. Ambient sound per location."
Varied reference angles give MVL a rounded model of the subject; explicitly licensing the outfit changes ('outfits changing per scene') stops the model from freezing clothing along with the face. The demo beside the generator is this prompt.
Start from: a clip you already generated
"Remove the passerby crossing behind the subject at the 4-second mark. Keep the subject, the camera drift, the lighting and all audio exactly as they are."
Timestamp the target, then protect everything else by name — the protection clause is what keeps an edit surgical. This is the workflow that makes 15-second renders cheap to perfect.
Upload: references for character A and character B
"A and B argue across a café table — A gestures with a spoon, B leans back laughing. Cut between over-shoulder shots. Each keeps their exact reference appearance. Espresso-machine hiss, low café chatter."
Naming who does what is coreference syntax: the model maps each action to the right referenced identity. Without it, two-character scenes trade faces mid-shot.
The reference-control triangle: three models that all promise 'your subject, kept consistent,' with different intake sizes and edit philosophies. Same prompts across all three, judged by what we'd ship. All three selectable in the generator above.
the O1 successor bringing the Omni architecture's native audio, multi-shot and conversational editing to two tiers. The version in the generator above.
Kuaishou's 'industry-first unified multimodal' bet: reference generation, in-painting, style transfer and shot extension fused into one engine. O3 is this idea, matured.
prompt-driven models that made Kling a global name, before the family split into the prompt-led V3 and reference-led O3 lines.
Same generation, different philosophy. Kling O3 (Video 3.0 Omni) is reference-first: your images and video define the subject, and MVL architecture keeps that identity locked across shots. Kling 3.0 (V3) is prompt-first cinematic work. Reference-driven briefs here, prompt-driven briefs on the Kling 3.0 page.