Describe the shot and pick the engine — thirteen models behind one prompt box, most of them generating sound alongside the picture. Three free drafts a day, no account.
Input
A prompt. Optionally a shot list
Models available
13, selectable per job
Clip length
5–30 s in one pass
Max resolution
Up to 4K on Kling 3.0
Audio
Native on most models
Free tier
3 drafts daily, no account
Typical time
30 s to 3 minutes
Our one-line verdict: the hardest thing AI video does, and the most fun. Write like a director rather than a describer, draft free until the shot reads, then spend credits once.
Text to video is the purest form of the tool and the least forgiving: with no photo to anchor composition, lighting and subject, the model invents all three from your sentence. Two things fix most disappointing results. First, prompt like a shot list — camera, subject, action, light, sound, in that order — because these models reward directorial language and ignore adjectives. Second, stop assuming one model does everything: a scene with cuts, a thirty-second unbroken take and a shot that must pass as real footage are three different engines. The table below routes you; the free tier means finding out costs nothing.
Thirteen engines, six jobs people actually arrive with. Each is selectable in the generator above — draft free first, then take the link.
The realism benchmark every summer flagship measured itself against, with ambience and spoken dialogue rendered alongside the picture. Eight seconds that survive scrutiny.
Director Mode composes up to six shots — angles and cuts included — inside one generation, at up to 4K. The only engine here that edits while it generates.
Thirty seconds in a single pass, with a supplied soundtrack able to drive pacing and lip-sync. Fifty reference files if the production has an asset kit.
Stop-motion, macro, nostalgic, deliberately strange — it escapes the glossy default look, and pins up to ten keyframes so the beats land where you put them.
Weight, momentum, pours and collisions rendered convincingly, with sequenced camera choreography and 21:9 framing. Silent output — plan an audio pass.
The only model here that reads a PDF, deck or spreadsheet and builds the video from its contents — thirty seconds, characters and layouts consistent with the source.

Camera, subject, action, light, sound — in that order. 'Medium shot, golden hour: a vendor slices a mango, juice catching the sun. Handheld, ambient market chatter.' Directorial beats descriptive every time.

Run the prompt on the free tier first. You're checking one thing: does the shot read? Composition and motion are visible even at draft quality, and iterating here costs nothing.

Take the routing table below, switch to the model that matches the job, and re-run the same prompt. Same words, better render — the prompt language transfers across models.
The right column is the part most tools leave out. Knowing the failure modes before you upload saves more time than any prompt guide.
• Nothing required but an idea.
No footage, no photo, no shoot. The distance between a thought and a moving image is one sentence and thirty seconds.
• Directorial control is real.
Camera moves, timing, lighting and sound genuinely respond to instruction on the stronger models — this isn't a slot machine if you prompt like a filmmaker.
• Sound arrives with the picture.
Most models here render ambience, effects and dialogue in the same pass, so a generated shot is publish-ready instead of needing a scoring session.
• Free drafting changes the economics.
Three free drafts a day means the expensive render happens once, on a shot you already know works.
• Lower hit rate than image-to-video.
With nothing anchoring composition, more attempts miss. If you have any usable still, the image-to-video page will get you there faster and cheaper.
• Text inside the video.
Signs, labels and captions come out unreliable on most video models — Kling 3.0 is the strongest here, but plan on adding critical text in your editor.
• Length and continuity.
Thirty seconds is the current ceiling for one unbroken take. Longer pieces are chained clips, and matching them takes editorial planning.
• Real people and public figures.
No generating recognizable real people, and prompts naming public figures are refused. Fictional characters, all yours.
The shot the stock library doesn't have, in the brand's own visual language, generated the afternoon it's needed rather than scheduled as a shoot.
Blocking a sequence before committing a crew day. Director Mode's multi-shot output makes a rough scene readable to people who don't read storyboards.
Historical scenes, scientific processes, anything a camera can't reach. Wan 3.0's document input turns an existing lesson plan straight into footage.
Write the prompt as a shot list — camera, subject, action, light, sound — draft it free to check the shot reads, then re-run it on whichever model the routing table above points to. Prompt language transfers between models, so the winning wording works everywhere.
Have a photo? Animating one gets a usable result faster.
Restyle or edit footage you already have.
Every engine, with one line each on what it wins.