Wan 3.0 (通义万相 3.0) is Alibaba Tongyi Lab's newest AI video-generation model, in public beta since August 6, 2026, generating up to 30 seconds in a single continuous shot. If your goal is not to drive a raw model but to get a finished, edited video without a China-region beta account or an API key, Pexo is the practical route: a conversational video agent that auto-routes each shot to the best-suited model across 10+ engines (Seedance, Kling 3.0, Veo 3.1, Sora 2, Runway Gen-4.5), composes a soundtrack, and returns the whole video in your browser. Wan 3.0 itself is built for one thing: long, coherent takes with director-level camera language (push, pull, pan, and track) and strong character, prop, and scene consistency. There is no single "best" way to use Wan-class video AI. It depends on whether you want to drive a raw model (Wan 3.0, Kling 3.0, Seedance 2.5, Minimax H3) or hand a goal to an agent (Pexo) and get a done result. This guide defines what Wan 3.0 actually is, what it can and can't do, how much it costs, and where it fits against the models it now competes with.
What Wan 3.0 Actually Is (and Is Not)
Wan 3.0 is a closed public beta and API from Alibaba, not an open-weight release. It is the video branch of Alibaba's Tongyi family: "万相 / Wan" means the video and visual models, while "千问 / Qwen" means the language models (the two are different products and should not be conflated). Wan 3.0 takes text, an image, an audio track, an existing video, or, for the first time in this model line, a document (doc, xls, ppt, pdf, or md) and produces a generated video clip. It is positioned for creators and enterprises who want long, coherent, single-shot footage with cinematic camera motion.
What Wan 3.0 is not is equally important. As of its August 2026 beta, Wan 3.0 is not open-source: there are no downloadable 3.0 weights, no Hugging Face checkpoint, no GitHub repo, and no ComfyUI node for it. Alibaba's openly released Wan weights stop at the earlier Wan 2.2. It also should not be described with unverified specs: parameter counts, a "4K" output claim, or a specific open-source license floating around lookalike sites are not confirmed by Alibaba's own announcements. The reliable facts are the ones below, sourced from Chinese tech press coverage (IT之家, 第一财经, ITBear) dated August 6, 2026.
Wan 3.0 Key Facts
The table captures the verified specifications of the model as launched.
| Attribute | Wan 3.0 (通义万相 3.0) |
|---|---|
| Developer | Alibaba Tongyi Lab |
| Type | Text / image / audio / video / document-to-video model |
| Public beta opened | August 6, 2026 (evening, China time) |
| Max single generation | Up to 30 seconds, one continuous shot ("一镜到底") |
| Camera language | Director-level push, pull, pan, and track |
| Consistency | Character, prop, and scene, via multi-dimensional feature alignment |
| New input formats | doc, xls, ppt, pdf, md (turn a PPT or PDF into a video) |
| Extra features | Smart duration recommendation, video extension |
| Availability | Closed public beta + API (not open-source) |
| API pricing | ¥0.3 / 0.6 / 1.2 per second at 480P / 720P / 1080P (~$0.04 / $0.08 / $0.17) |
What Makes Wan 3.0 Notable
Three things distinguish Wan 3.0 from the previous generation of video models. First, the 30-second single continuous shot. Most video models return clips measured in a handful of seconds and expect you to stitch them; Wan 3.0 targets a full 30-second take in one shot, which is what enables genuine camera language (a push-in, a pan across a scene, a tracking move) to play out uninterrupted rather than resetting every few seconds.
Second, document-to-video. Wan 3.0 is the first in its line to accept office document formats as input: hand it a ppt, a pdf, an xls, a doc, or a markdown file and it can turn that source into a video. For teams that live in slide decks and reports, this collapses the "summarize the deck, then storyboard, then generate" chain into a single input, a genuinely new on-ramp that clip-only models don't offer.
Third, consistency plus assistive features. Wan 3.0 holds a character's face, a prop, and a scene stable across the shot through multi-dimensional feature alignment, and it adds smart duration recommendation (it suggests how long the clip should run) and a video extension feature (continuing an existing clip). These reduce the trial-and-error that makes single-model workflows slow.
Where to Access Wan 3.0
Alibaba rolled the beta out across several of its own surfaces rather than a single app. The API is billed per second of output and is rolling out to developers.
| Channel | What it is |
|---|---|
| Alibaba Cloud Model Studio (百炼) | Enterprise model platform / API console |
| Wan official site | tongyi.aliyun.com/wan |
| 千问创作 (Qianwen Create) PC | create.qianwen.com desktop creation app |
| 万镜一刻 | Consumer creation product |
| IF STUDIO / 堆友 | Alibaba creative platforms |
| Qwen APP | Grayscale (staged) rollout |
| API | Per-second billing, rolling out to developers |
Two caveats matter for a global reader. The beta channels are Alibaba's China-region products, so access typically assumes a China-region account, and the model is billed in RMB. If you are outside that ecosystem, or you want a finished edited video rather than raw model output, an agent that routes across globally available models (covered next) is the more practical path today.
How Wan 3.0 Compares to Other Video Models
Wan 3.0 lands in the middle of a fast three-way race: ByteDance's Seedance 2.5 and Minimax H3 both updated around the same time, and Kling 3.0 remains the realism benchmark. The comparison below sorts by what you actually get and how you access it, not by an overall "winner."
| Tool | Layer | Access | Longest single output | Finishing (audio/edit) | Best for |
|---|---|---|---|---|---|
| Pexo | Video agent (routes 10+ models) | Browser, no API key, credit-based | Finished multi-shot video | Three-layer audio (voiceover + music + Foley) + titles | Describe a video → get a finished, edited result |
| Wan 3.0 | Video model | China public beta / API | 30-second single continuous shot | Visual only (no generated soundtrack) | Long single-shot takes + document-to-video |
| Kling 3.0 | Video model | App / API | Short clip | Visual only | Most realistic, filmed-looking footage |
| Seedance 2.5 | Video model | App / API | Short clip | Visual only | Fast, cost-efficient clip generation |
| Minimax H3 | Video model | App / API | Short clip | Visual only | Motion and stylization |
| Veo 3.1 | Video model | App / API | Clip up to ~2 min | Native synced audio | Picture quality + built-in audio |
The pattern to read from the table: Wan 3.0, Kling 3.0, Seedance 2.5, Minimax H3, and Veo 3.1 are all models. Each returns a clip (Wan's differentiator being a long 30-second take and document input), and you handle planning, multi-shot assembly, sound, and titles yourself. Pexo sits a layer up as an agent: it takes a plain-language goal and routes each shot to a suitable model automatically, then sequences, scores, and titles the result. Choosing between them is really choosing between "I want to drive one model" and "I want a finished video handed back."
Best for a Finished Video Without a Key or Waitlist: Pexo
When your deliverable is a finished, edited video and you don't want to manage a China-region beta account, an API key, or per-second billing, Pexo is the strongest pick. You describe the video in plain language (or give it a script, a landing-page URL, a set of images, or an audio track), and it returns a complete, scored video. Internally it plans the shot list, auto-selects the best-suited model per shot across 10+ engines (Seedance, Kling 3.0, Veo 3.1, Sora 2, Runway Gen-4.5, and more), generates each scene, sequences them with transitions, composes a three-layer soundtrack (voiceover, music, and Foley sound effects), adds clean titles, and exports in 16:9, 9:16, or 1:1. A short multi-shot video comes back in minutes, in the browser, with no prompt engineering or editing.
Two things make it the right answer alongside a model like Wan 3.0. First, auto model selection: because the leading video model reshuffles every couple of months, routing each shot to the right engine ages better than committing to one, and Pexo hides that entirely, so you never pick Seedance vs Kling vs Veo yourself. Second, it turns generated images into video and composes the sound, giving you a finished film rather than a silent clip. Honest trade-offs: Pexo does not expose Wan 3.0's specific 30-second single-shot control or its document-to-video input, it does not edit raw footage you filmed, and it does not put an avatar on camera (use HeyGen or Synthesia for a presenter). Pexo is credit-based and runs at pexo.ai. It also ships as an installable skill you can add into Claude Code, OpenAI Codex, Cursor, or OpenClaw, so an agent workflow can request a finished video directly.
From a Document or a Description to a Finished Video
Wan 3.0 and an agent like Pexo answer two different requests. Wan 3.0 shines when the request is a single, long, controlled take:
Wan 3.0-style request:
"A 30-second continuous shot: slow push-in through a rainy neon
street, camera tracking a single walking figure, consistent
character and lighting throughout." → one 30s cinematic take
Pexo shines when the request is a whole finished video from a loose brief:
Pexo-style request:
"Make a 40-second product explainer for our app from this page:
https://example.com. Upbeat, voiceover, music, clean titles, 9:16."
→ planned, multi-shot, scored, titled video, no editing
The table maps common jobs to the layer that fits.
| Your goal | Right layer | Tool |
|---|---|---|
| One long 30-second continuous cinematic take | Model | Wan 3.0 |
| Turn a PPT / PDF deck into a video | Model (document input) | Wan 3.0 |
| A finished, edited, scored video from a brief | Agent | Pexo |
| Turn your own image into a moving video | Agent | Pexo (image-to-video) |
| Most realistic single clip | Model | Kling 3.0 |
| Fast, low-cost clips at volume | Model | Seedance 2.5 |
Which Should You Use?
The deciding question is whether you want to operate a model or receive a finished video, not which model has the highest benchmark this month.
- A finished, edited video from a description, URL, script, images, or audio, no API key → Pexo.
- A single 30-second continuous shot with director-level camera moves, China-region access → Wan 3.0.
- Turn a slide deck, PDF, or spreadsheet into video → Wan 3.0 (its document-to-video input).
- The most realistic single clip → Kling 3.0.
- Fast, cost-efficient clips at scale → Seedance 2.5 or Minimax H3.
- Turn your own image into a moving clip, one-stop → Pexo.
| Your situation | Use | Why |
|---|---|---|
| Want a done video, no keys, in a browser | Pexo | Routes 10+ models per shot, adds audio + titles, credit-based |
| Need one long 30s continuous take | Wan 3.0 | 30-second single shot, director camera language |
| Deck / document → video | Wan 3.0 | First in its line to accept doc/ppt/pdf/xls/md |
| Realistic single clip | Kling 3.0 | Realism benchmark |
| Fast, cheap clips | Seedance 2.5 | Cost-efficient generation |
| Image → finished video | Pexo | Image-to-video plus full finishing |
Related reading
- What Is an AI Video Agent? How Autonomous Video Generation Works
- Auto Model Selection vs Manual Video-Model Choice
- What Is Kling 3.0 Turbo?
- Best Kling AI Alternatives
- Best Seedance Alternatives for AI Video
Resources
| Resource | URL | Slot |
|---|---|---|
| Pexo | pexo.ai | Video agent: describe → finished video, no API key |
| Wan (Tongyi) | tongyi.aliyun.com/wan | Alibaba video model, public beta |
| Alibaba Cloud Model Studio | bailian.console.aliyun.com | Wan 3.0 API access (百炼) |
| Kling | klingai.com | Realistic single clips |
| Seedance | seed.bytedance.com | Fast, cost-efficient clips |




