Wan 3.0 AI Video Generator

Alibaba's new all-in-one model: native 30-second clips from a prompt, an image — or a PDF, slide deck or spreadsheet. The only model here that reads documents. Type a prompt below and try it, no waitlist.

Model
Upload Image
Click or drop an image here JPG, JPEG, PNG or WEBP up to 10 MB
Prompt
0/5000
Output Aspect Ratios
Auto
Resolution
Auto
Duration(seconds)
Auto
Video Sample
 

Developer

Alibaba Tongyi Lab

Type

All-in-one: gen + ref + edit

Duration

2–30 s single pass

Max output

1080p

Doc input

pdf / ppt / xls / URL

Open weights

Pledged (Apache 2.0)

Tier here

Paid credits

Our one-line verdict: the document pick. When the source material is a deck, a manual or a spreadsheet, nothing else can even read it.

Wan 3.0 went into public beta on August 6th, and its sharpest edge isn't the 30-second single take — Seedance got there first — it's what it accepts as input. Hand it a PDF manual, a pitch deck or a quarterly spreadsheet and it builds the video from the contents, keeping characters, props and layouts consistent with the source. Alibaba also collapsed what used to be four separate Wan models into one, so generation, reference and editing no longer mean switching models mid-pipeline.

Honest read

Where Wan 3.0 wins — and where it doesn't

Public beta is days old, so these are first impressions from our test renders plus Alibaba's published documentation — flagged as such. We'll revise as independent numbers land.

Genuinely strong

  • • Document-to-video.

    Reads doc, xls, ppt, pdf and similar formats up to 100 MB or 50 pages, plus web pages by URL, and generates video from the contents. A training video from a manual, a demo from a deck — no other model we host can do this at all.

  • • The 30-second single take.

    One native generation from 2 to 30 seconds, double the previous Wan ceiling, with an intelligent-duration feature that suggests the right length for your prompt.

  • • All-in-one workflow.

    Wan 2.7 split text-to-video, image-to-video, reference and editing across four models. Wan 3.0 merges them — one model ID, no pipeline switching, references and edits in the same place.

  • • Reference fidelity.

    Characters, props, voices, spatial layouts and style stay consistent with your inputs — Alibaba's framing is "from generating a frame to telling a whole story," and our early tests back the continuity claim.

Known limits

  • • 1080p ceiling.

    Documented outputs are 480P, 720P and 1080P — no 2K tier. If resolution is the brief, MiniMax H3 holds that card.

  • • Access friction upstream.

    On Alibaba Cloud the model is labeled preview and pricing is invitation-only. Here you skip the waitlist, but expect upstream capacity hiccups while the beta scales.

  • • Open weights: pledged, not shipped.

    Apache 2.0 is promised; the Wan line's record on delivering open releases has been uneven. Believe the upload, not the announcement.

  • • Benchmarks pending.

    It launched into the most crowded month in AI video history — every "best model" claim about it, its rivals, or by us, is provisional until independent scores land.

Prompt recipes

Three prompts that show what it's for

Two of these exercise the document trick nobody else has; the third is the long take. Copy, swap the brackets, render.

The deck-to-launch-video

Upload: your product deck (ppt/pdf, ≤50 pages)

"Turn this deck into a 30-second launch video: a presenter walks through the three key features in order, product renders match the slides exactly, upbeat corporate tone, end on the logo and tagline from the final slide."

Why it works

Naming the structure ("three key features in order", "end on the final slide") tells the model how to spend its 30 seconds on your 12 slides. "Match the slides exactly" invokes the reference-fidelity system — it's reading your deck, not decorating around it.

The manual-to-training-clip

Upload: an equipment manual (pdf)

"A 30-second safety training clip from section 3 of this manual: an operator demonstrates the startup sequence step by step, close-ups on each control named in the text, calm instructional voiceover, captions matching the manual's terminology."

Why it works

Pointing at a section keeps a 50-page input focused, and "terminology matching the manual" is the instruction that makes the output usable for actual training instead of generic b-roll. Spreadsheet version: swap in an xls and ask for a chart-driven quarterly summary.

The unbroken story shot

"One continuous 30-second shot: a [street food vendor] sets up her stall at dawn — unfolds the cart, lights the burner, first customer arrives as the sun clears the rooftops. The camera slowly circles the stall the whole time. Ambient morning sounds building to sizzling and chatter."

Why it works

A beginning-middle-end written as one sentence of continuous action is the shape long-take models reward. The circling camera is the continuity stress test — if character and layout hold through 360 degrees for 30 seconds, the model is doing its job.

Head to head

Wan 3.0 vs Seedance 2.5 vs MiniMax H3

Three Chinese labs, three flagships, six weeks. Same prompts across all three, judged by what we'd ship — provisionally, given how new all of them are. Compare them yourself from one prompt box.

JobWan 3.0Seedance 2.5MiniMax H3
Document input (pdf/ppt/xls) Yes — unique No No
Max clip length 30 s native 30 s native (180 s beta) 15 s (extend ~30 s)
Max resolution 1080p High-res via family pipeline 2K native
Audio-driven pacing from a track No Yes Voice reference
All-in-one (gen + ref + edit) One model Region edits, separate flows Instruction edits
Cost per clip Medium Highest (long renders) Lowest for 2K
Open-weight roadmap Apache 2.0 pledged None Pledged, unshipped
Access without waitlist Invite-only upstream Open Open
Version history

The Wan line, briefly

Aug 2026 · Current

Wan 3.0

public beta August 6 as wan3.0-video on Alibaba Cloud. Native 30-second clips, document-to-video, all-in-one architecture, Apache 2.0 open weights pledged. The version on this page.

Earlier

Wan 2.6 / 2.7

the commercial era: strong 15-second generation with the workflow split across four specialized models, and an open-source story that ran hot and cold. The versions most existing Wan reviews describe.

Origin

Wan 2.1 / 2.2 — the open era

the releases that built the Wan community: capable open weights that spawned fine-tunes and local workflows, and the reason "will the weights actually ship" is the first question anyone asks about Wan 3.0.

FAQ

Wan 3.0 questions, answered straight

1. What is Wan 3.0?

The new flagship of Alibaba Tongyi Lab's Wan video line, launched in public beta on August 6, 2026 under the model ID wan3.0-video. Three headline changes: native 30-second single-pass clips (double Wan 2.7's ceiling), an omni-reference system that accepts documents — PDFs, decks, spreadsheets, even web pages — as creative inputs, and an all-in-one architecture merging what used to be four separate Wan models into one.

Other models

Not the right tool? Try its neighbors

Seedance 2.5

The other long-take model. 30-second shots cut to your soundtrack.

MiniMax H3

The 2K pick. Cheapest high-res clips with locked characters and voices.

FLUX 3

The stylist. 20-second clips with audio, from cinematic to stop-motion.