Black Forest Labs' first video model: 20-second clips with synchronized audio, up to 10 pinned keyframes, and looks that run from live-action to stop-motion. Type a prompt below and try it.
Developer
Black Forest Labs
Type
Multimodal (video now, image soon)
Audio
Native, synced, same pass
Duration
Up to 20 s per pass
Keyframes
Up to 10 pinned + first/last
Open weights
Dev planned, late 2026
Tier here
Paid credits + draft mode
Our one-line verdict: the storyboard pick. When you know exactly which frames the shot must hit — or you want a look that isn't "cinematic AI video" — start here.
FLUX 3 comes from the Freiburg lab whose founders built the architecture behind Stable Diffusion, and it's their first model that moves: one network jointly trained on images, video, audio and action, shipping video first. Two things set it apart in a crowded month — you can pin up to ten exact frames and let the model animate between them, and the output isn't locked to the glossy "AI cinematic" look. Raw, nostalgic, handmade, strange: the range is the point.
Early impressions from our test renders plus BFL's published material — with their own caveats passed along intact. We'll revise as independent numbers land.
• Keyframe control.
Pin up to 10 exact frames plus first-and-last, and the model generates motion, camera, dialogue and audio between them. Storyboarding as an input format — nothing else we host accepts this many anchors.
• Aesthetic range.
Stop-motion, macro, time-lapse, nostalgic home-video, deliberately strange — it escapes the uniform cinematic gloss most video models default to. If your brand look isn't "movie trailer," this matters.
• Audio in the same pass.
Up to 20 seconds with a synchronized soundtrack — dialogue included — and seam-aware extension that carries movement, camera and audio across the join when you extend a clip.
• Draft mode economics.
Every endpoint has a fast draft variant for cheap previews. Iterate on the draft, spend real credits once on the final — the workflow this whole site preaches, built into the model family.
• The image part isn't here yet.
The FLUX name is famous for images, but FLUX 3's image generation is still rolling out. If you came here to make pictures, that's weeks away — see the FAQ before you burn credits finding out.
• Benchmarks are house numbers.
The launch chart was labeled a preliminary evaluation of an early candidate, and the flashiest win rates came against models that aren't setting the current pace. BFL's later claims are stronger but still internal.
• 20 seconds, not 30.
Wan 3.0 and Seedance 2.5 both reach 30 in one pass. FLUX 3 answers with extension that respects the seam — but if a single unbroken half-minute is the brief, it's not the pick.
• Dev weights: planned, not dated.
The open-weight release is announced for later in 2026 with no date or license terms. BFL's open-source record is genuinely good, but a plan is still a plan.
One for the look, one for the keyframes, one for the seam. Copy, swap the brackets, render — in draft mode first.
"A handmade stop-motion scene: a clay [astronaut] plants a tiny paper flag on a kitchen-table moon made of flour. Visible fingerprints in the clay, slightly jerky 12fps motion, soft window light, felt-footstep sounds and a muffled line of dialogue in a squeaky voice."
Naming the imperfections — fingerprints, jerky frame rate — is what steers FLUX 3 away from the default gloss. Most models fight you on "handmade"; this one treats it as a style target.
Upload: 4 keyframes (product boxed → unboxed → in hand → logo card)
"A 20-second unboxing that hits these four frames in order, roughly 5 seconds apart. Natural handheld feel, paper-tearing and tape sounds, quiet delight from an off-screen voice at frame three."
The frames carry the composition so the prompt only has to direct timing, motion and sound. Spacing hints ("roughly 5 seconds apart") stop the model from rushing the beats. This is the workflow storyboard-first teams have been waiting for.
Start from: a clip you already generated
"Extend by 10 seconds: the [dancer] completes the turn the clip ends on, the camera keeps drifting left at the same speed, the music continues without a hitch — no cut, no reset."
Extension reads the final frames' motion and audio, so the instruction should describe continuation, not a new scene. "Completes the turn the clip ends on" hands it the physical thread to pull; naming the camera speed protects the drift across the seam.
Three audio-native models, three control philosophies. Same prompts across all three, judged by what we'd ship — provisionally for the newer two. Compare them yourself from one prompt box.
announced July 23: one architecture for image, video, audio and action. Video shipped first (20 s, native audio, keyframes); image generation rolling out in the following weeks; the open-weight Dev release planned for later in the year. The version on this page.
the image era that made the name: FLUX.1 dev became one of the most downloaded and fine-tuned open diffusion releases anywhere, the backbone of countless community workflows.
the founding team built the latent diffusion architecture behind Stable Diffusion. The open-source credibility that history buys is why "FLUX 3 Dev" carries more weight as a promise than most open-weight pledges do.
Black Forest Labs' first multimodal foundation model, announced July 23, 2026 — one architecture jointly trained on images, video, audio and action prediction, from the Freiburg team whose founders built the architecture behind Stable Diffusion. The part you can generate with today is video: up to 20 seconds with a synchronized soundtrack in the same pass, from text, a starting image, pinned keyframes, or an existing clip you want to extend.
The realist. When the clip has to pass as actual footage, with audio.
The long take. 30-second shots cut to your soundtrack.
The document reader. Decks, PDFs and spreadsheets in, video out.