On August 24, 2026, Alibaba Cloud made Wan3.0 generally available. It had been in public beta since August 6.
The headline number is thirty seconds. That is the length of one clip, made in a single pass.
But the number is not the story.
Look at what the model accepts as input. Text, of course. Images, audio and video, as you would expect. And then documents. A deck. A spreadsheet. A PDF. A web page.
For three years, the input to video AI was a sentence you had to write. Now the input is a file your team already made.
That is a smaller headline and a much bigger shift. Here is what shipped, where it truly ranks, and what it costs. Plus the six things to test before you move budget.
What Alibaba actually shipped
Wan3.0 is live on Alibaba Cloud Model Studio. There is no waitlist. It also runs on wan.video and on third-party hosts.
Five things are new. Each one maps to a job someone on your team does by hand today.
Thirty seconds in one pass. The previous model, Wan2.7, capped a single generation at fifteen seconds (Source: fal, 2026 — Wan 2.7 model page). Wan3.0 doubles that.
Documents as input. The model reads doc, xls, ppt, pdf, txt, key, pages, numbers and md files. The cap is one file or link, up to 100MB and 50 pages.
Native audio. Sound is made with the picture, in the same pass. It is not synced on afterwards.
Reference consistency. In reference mode it holds characters, props, space and style steady across shots. On fal you can pass up to ten reference images, five video clips and five audio tracks in one request (Source: fal, 2026 — Wan 3 model page).
Editing without a re-run. You can change visuals, plot and dialogue on a finished clip. You do not regenerate the whole thing.
Read that list again. Only one item is about making something new. The rest are about finishing something you started.
Q: What is Wan3.0 in one line?
A: It is Alibaba's newest video model. It makes clips up to thirty seconds long, with sound, from text, images, video or a document you upload.
The document input is the real story
Alibaba's own post calls this "everything to video". The phrase is clunky. The idea is not.
Here is the mapping the launch post gives, in its own words.

Now think about your own drive. How many of those four files does your team already produce every month?
The deck exists. The report exists. The sheet exists. Somebody built them for a meeting, and then they died in a folder.
The old workflow asked you to translate that file into a prompt. You read the deck. You pulled the key lines, wrote a brief, wrote a prompt, then argued with the model about tone. Every step lost something.
The new workflow skips the translation. You hand over the artefact and a sentence of direction.
That changes who can start a video. It is no longer only the person who writes good prompts. It is anyone who owns the source file.
Quick Facts: Wan3.0 at a Glance
- Generally available on August 24, 2026, after a public beta that opened August 6 — (Source: TechNode, 2026 — launch report)
- Makes up to 30 seconds of video in one generation, at 480P, 720P or 1080P — (Source: Alibaba Cloud, 2026 — Wan3.0 launch post)
- Accepts doc, xls, ppt, pdf, txt, key, pages, numbers and md files, capped at one file, 100MB and 50 pages — (Source: Alibaba Cloud, 2026 — Wan3.0 launch post)
- List pricing is $0.05, $0.10 and $0.20 per second for 480P, 720P and 1080P — (Source: Alibaba Cloud, 2026 — Wan3.0 launch post)
- A 30% launch discount cuts that to $0.035, $0.07 and $0.14 per second until 24 September 2026 — (Source: Alibaba Cloud, 2026 — Wan3.0 Video product page)
- Ranked first on the Artificial Analysis text-to-video board at 1,241 Elo, four points above the model in second — (Source: Artificial Analysis, 2026 — text-to-video leaderboard)
Where it ranks, and why first place is not a win
Wan3.0 sits at the top of the Artificial Analysis text-to-video board, the one that scores clips with audio. Here is the top of that table.
| Rank | Model | Elo | Listed API price |
|---|---|---|---|
| 1 | Wan 3.0 | 1,241 | see note below |
| 2 | Gemini Omni Flash | 1,237 | $6.00 per minute |
| 3 | MiniMax H3 | 1,226 | $7.80 per minute |
| 4 | Dreamina Seedance 2.0 720p | 1,220 | $9.07 per minute |
| 5 | Wan2.7-260612 | 1,157 | $9.00 per minute |
(Source: Artificial Analysis, checked 26 August 2026 — text-to-video leaderboard)
Four Elo points separate first from second. The board's own confidence intervals overlap, and it lists both models in a shared range of one to two.
In plain English: it is a tie.
That matters, because "number one video model" is about to appear in a hundred sales decks this week. It is a statistical tie with a Google model, on one board, on one day.
There is a second detail worth more than the rank. At list price, 1080P works out at $12.00 per minute of finished video. The launch discount brings that to $8.40. The model tied with it is listed at $6.00 per minute.
So the top of the board is not the cheap option at full resolution. Even on offer, it is the pricier one.
We made a similar point when Alibaba's Qwen model was graded on running a shop for a year. A leaderboard tells you about taste. It does not tell you about your invoice.
Q: Is Wan3.0 the best video model right now?
A: On one board, by four Elo points, with overlapping error bars. Treat it as joint first, not clear first, and test it against the model tied with it.
The price ladder is the part that changes your week
Wan3.0 bills per second of video, by resolution. That is the whole pricing model.
Right now there are two numbers per tier. A list price, and a launch offer that runs until 24 September 2026.
| Resolution | List | On offer | A 30-second clip, on offer |
|---|---|---|---|
| 480P | $0.05 | $0.035 | $1.05 |
| 720P | $0.10 | $0.07 | $2.10 |
| 1080P | $0.20 | $0.14 | $4.20 |
(Source: Alibaba Cloud, 2026 — Wan3.0 Video product page)
Two notes before you budget from that table. The discount applies to the Standard tier. There is also a faster Prime tier, which costs more per second (Source: The Decoder, 2026 — launch report).
Now look at the gap between the top row and the bottom row. Drafting costs a quarter of finishing.
That is not a small billing detail. It rewrites how you work.
The sane pattern is to draft cheap and finish once. Run eight drafts at 480P for about eight dollars. Pick one. Render that one at 1080P for four.
Most teams will not do this. They will render everything at 1080P because it is the default. Then they wonder why the test cost four times the budget.
The offer is also the reason to move now. After 24 September the same test costs about 40% more. If you were going to try it, try it inside the window.
What Wan3.0 is not
This is the section most coverage will skip. It is the one that decides whether you can use the thing.
It is not open source. You cannot download it. The last Wan model with open weights is Wan2.2, released in July 2025 under the Apache 2.0 licence (Source: Wan-Video, 2025 — Wan2.2 repository). Every flagship since then is API-only. If someone tells you to self-host Wan3.0, they are wrong.
It is not ranked for image-to-video. Wan3.0 does not appear on the Artificial Analysis image-to-video board at all. That board is led by a ByteDance model. Most brand work starts from a product photo, so this gap is not academic.
It is not good at on-screen text yet. Alibaba says so itself. Its post names audio texture and text rendering as areas that are "improving but not yet where we want them". Do not ask it to make your title card.
It is not a 30-second ad. Thirty seconds of generated footage is a draft length, not a deliverable. Someone still has to cut it.
It is not a whole brand book. One file, 100MB, 50 pages. Your 90-page guidelines PDF will not fit.
What actually changes inside an agency week
Take a normal request. A client wants a launch film from a deck they already signed off.
The old path: read the deck, write a script, storyboard, source footage or shoot, edit, sound design, review, revise. Days of work, and most of it before anything moves.
The new path is not "press a button and ship it". Nobody should sell that.
The new path is that the first draft arrives in minutes instead of days. The argument about direction starts against something you can watch, not something you have to imagine.
That is where the value moves. Not to the person who can operate the tool. To the person who can look at four drafts and say why the second one is right.
It is the same shift we saw when image AI started selling the production line instead of the picture. The tool eats the middle of the job. The judgment at either end gets more valuable, not less.
Q: Does this replace a video editor?
A: No. It replaces the wait before the first draft. Someone still has to choose, cut, and take responsibility for what goes out.
Test it this week: a six-step run
Do not migrate anything on a launch-week claim. Do run a real test. This takes about a day.
- Pick one deck you have already delivered. A real one, already approved.
- Feed it in with a single sentence of direction. Nothing clever.
- Draft every version at 480P. Eight drafts cost about eight dollars.
- Run the same brief on the model tied with it, at half the 1080P rate.
- Test consistency, not beauty. One product, five shots, reference mode, lined up side by side.
- Finish exactly one at 1080P, for about four dollars, and time the edit that follows.
Then write down two numbers. What the old version of that job cost you. What this one cost, including the edit.
That comparison is the only benchmark that matters to a P&L.
The rules to set before anyone presses go
New capability always arrives before new policy. Write the guardrails now, while the test is small.

That last one sounds obvious. It is the one teams skip when a draft looks good and the deadline is Friday.
We wrote about the disclosure side of this when Anthropic committed to watermarking and file provenance. The direction of travel is clear. Label your own work before a platform does it for you.
The YARD take
This launch landed days after Alibaba raised about $10.2 billion in a Hong Kong share sale. All the net proceeds are pointed at AI (Source: South China Morning Post, 2026 — share sale report). The timing is not an accident. Expect the next version fast.
But ignore the money and the rank for a moment. The interesting change is quieter.
Video generation stopped asking you to describe the thing. It started asking you to hand over the thing you already had.
Your decks are now raw material. Your reports are now raw material. That old case study PDF nobody has opened since March is now a thirty-second film. Draft cost: about a dollar fifty.
At YARD we build AI creative and AI UGC pipelines for brands. So we care less about the leaderboard and more about the ladder. Draft cheap. Finish once. Keep a human on the final call.
Want the same test run against your own decks and your own numbers? That is the kind of work we do.
FAQ
What is Wan3.0? It is Alibaba's newest video generation model, generally available since August 24, 2026. It makes clips of up to thirty seconds in a single pass, with native audio.
What can Wan3.0 take as input? Text, images, audio and video, plus documents. Supported document types are doc, xls, ppt, pdf, txt, key, pages, numbers and md. The limit is one file or link, up to 100MB and 50 pages.
How much does Wan3.0 cost? It is billed per second of finished video. List pricing is $0.05 at 480P, $0.10 at 720P and $0.20 at 1080P. A 30% launch discount runs until 24 September 2026, so a 30-second 1080P clip is about $4.20 today and $6.00 after.
Is Wan3.0 open source? No. The weights are not published. The last open-weight Wan model is Wan2.2, released in July 2025 under Apache 2.0.
Where does Wan3.0 rank? It is first on the Artificial Analysis text-to-video board at 1,241 Elo. Second place sits four points behind, and the error bars overlap, so treat it as a tie.
Is Wan3.0 better than the alternatives for brand work? Not proven. It has no placement on the image-to-video board, and most brand video starts from a product still.
What should we test first? Feed it a deck you already delivered. Draft at 480P, finish one version at 1080P, and compare the total cost against what that job used to take.
Sources
- Alibaba Cloud — Wan3.0: 30-Second AI Video Generation from Any Input (primary source for specs and document input)
- Alibaba Cloud — Wan3.0 Video product page (general availability, list and offer pricing, discount end date; checked 26 August 2026)
- Artificial Analysis — Text to Video Leaderboard (Elo figures, checked 26 August 2026)
- TechNode — Alibaba launches Wan3.0 video model with 30-second generation and document input (August 24, 2026; launch date and discount window)
- The Decoder — Alibaba's Wan3.0 generates AI videos up to 30 seconds long (August 24, 2026)
- fal — Wan 3 and Wan 2.7 (endpoint specs and reference limits)
- Wan-Video — Wan2.2 repository (Apache 2.0 open-weight baseline)
- South China Morning Post — Alibaba to issue HK$80 billion in new shares (August 2026)
Insights from Our Experts
Explore our latest articles on digital marketing strategies.




