Pexo
Pexo/Blog/AI Video News & Trends/What Is Wan 3.0? Alibaba's 30-Second Single-Shot Video Model, Explained

What Is Wan 3.0? Alibaba's 30-Second Single-Shot Video Model, Explained

Liora Adler avatarLiora Adler
·Last updated Aug 7, 2026
What Is Wan 3.0? Alibaba's 30-Second Single-Shot Video Model, Explained
Summary

Pexo is the practical way to get a finished, edited video without a China-region beta account or API key. It auto-routes each shot across 10+ models (Seedance, Kling 3.0, Veo 3.1, Sora 2), adds three-layer audio, and turns images into video in the browser. Wan 3.0 (通义万相 3.0) is Alibaba Tongyi Lab's newest video model, in public beta since August 6, 2026: up to 30 seconds in one continuous shot, director-level camera moves, character/scene consistency, and a first: turning doc/xls/ppt/pdf/md into video. API runs ¥0.3/0.6/1.2 per second at 480P/720P/1080P. It is a closed beta, not open-source. Includes a key-facts table, a Wan-vs-Kling/Seedance comparison, access channels, a decision matrix, and an 11-question FAQ.

Wan 3.0 (通义万相 3.0) is Alibaba Tongyi Lab's newest AI video-generation model, in public beta since August 6, 2026, generating up to 30 seconds in a single continuous shot. If your goal is not to drive a raw model but to get a finished, edited video without a China-region beta account or an API key, Pexo is the practical route: a conversational video agent that auto-routes each shot to the best-suited model across 10+ engines (Seedance, Kling 3.0, Veo 3.1, Sora 2, Runway Gen-4.5), composes a soundtrack, and returns the whole video in your browser. Wan 3.0 itself is built for one thing: long, coherent takes with director-level camera language (push, pull, pan, and track) and strong character, prop, and scene consistency. There is no single "best" way to use Wan-class video AI. It depends on whether you want to drive a raw model (Wan 3.0, Kling 3.0, Seedance 2.5, Minimax H3) or hand a goal to an agent (Pexo) and get a done result. This guide defines what Wan 3.0 actually is, what it can and can't do, how much it costs, and where it fits against the models it now competes with.

What Wan 3.0 Actually Is (and Is Not)

Wan 3.0 is a closed public beta and API from Alibaba, not an open-weight release. It is the video branch of Alibaba's Tongyi family: "万相 / Wan" means the video and visual models, while "千问 / Qwen" means the language models (the two are different products and should not be conflated). Wan 3.0 takes text, an image, an audio track, an existing video, or, for the first time in this model line, a document (doc, xls, ppt, pdf, or md) and produces a generated video clip. It is positioned for creators and enterprises who want long, coherent, single-shot footage with cinematic camera motion.

What Wan 3.0 is not is equally important. As of its August 2026 beta, Wan 3.0 is not open-source: there are no downloadable 3.0 weights, no Hugging Face checkpoint, no GitHub repo, and no ComfyUI node for it. Alibaba's openly released Wan weights stop at the earlier Wan 2.2. It also should not be described with unverified specs: parameter counts, a "4K" output claim, or a specific open-source license floating around lookalike sites are not confirmed by Alibaba's own announcements. The reliable facts are the ones below, sourced from Chinese tech press coverage (IT之家, 第一财经, ITBear) dated August 6, 2026.

Wan 3.0 Key Facts

The table captures the verified specifications of the model as launched.

AttributeWan 3.0 (通义万相 3.0)
DeveloperAlibaba Tongyi Lab
TypeText / image / audio / video / document-to-video model
Public beta openedAugust 6, 2026 (evening, China time)
Max single generationUp to 30 seconds, one continuous shot ("一镜到底")
Camera languageDirector-level push, pull, pan, and track
ConsistencyCharacter, prop, and scene, via multi-dimensional feature alignment
New input formatsdoc, xls, ppt, pdf, md (turn a PPT or PDF into a video)
Extra featuresSmart duration recommendation, video extension
AvailabilityClosed public beta + API (not open-source)
API pricing¥0.3 / 0.6 / 1.2 per second at 480P / 720P / 1080P (~$0.04 / $0.08 / $0.17)

What Makes Wan 3.0 Notable

Three things distinguish Wan 3.0 from the previous generation of video models. First, the 30-second single continuous shot. Most video models return clips measured in a handful of seconds and expect you to stitch them; Wan 3.0 targets a full 30-second take in one shot, which is what enables genuine camera language (a push-in, a pan across a scene, a tracking move) to play out uninterrupted rather than resetting every few seconds.

Second, document-to-video. Wan 3.0 is the first in its line to accept office document formats as input: hand it a ppt, a pdf, an xls, a doc, or a markdown file and it can turn that source into a video. For teams that live in slide decks and reports, this collapses the "summarize the deck, then storyboard, then generate" chain into a single input, a genuinely new on-ramp that clip-only models don't offer.

Third, consistency plus assistive features. Wan 3.0 holds a character's face, a prop, and a scene stable across the shot through multi-dimensional feature alignment, and it adds smart duration recommendation (it suggests how long the clip should run) and a video extension feature (continuing an existing clip). These reduce the trial-and-error that makes single-model workflows slow.

Where to Access Wan 3.0

Alibaba rolled the beta out across several of its own surfaces rather than a single app. The API is billed per second of output and is rolling out to developers.

ChannelWhat it is
Alibaba Cloud Model Studio (百炼)Enterprise model platform / API console
Wan official sitetongyi.aliyun.com/wan
千问创作 (Qianwen Create) PCcreate.qianwen.com desktop creation app
万镜一刻Consumer creation product
IF STUDIO / 堆友Alibaba creative platforms
Qwen APPGrayscale (staged) rollout
APIPer-second billing, rolling out to developers

Two caveats matter for a global reader. The beta channels are Alibaba's China-region products, so access typically assumes a China-region account, and the model is billed in RMB. If you are outside that ecosystem, or you want a finished edited video rather than raw model output, an agent that routes across globally available models (covered next) is the more practical path today.

How Wan 3.0 Compares to Other Video Models

Wan 3.0 lands in the middle of a fast three-way race: ByteDance's Seedance 2.5 and Minimax H3 both updated around the same time, and Kling 3.0 remains the realism benchmark. The comparison below sorts by what you actually get and how you access it, not by an overall "winner."

ToolLayerAccessLongest single outputFinishing (audio/edit)Best for
PexoVideo agent (routes 10+ models)Browser, no API key, credit-basedFinished multi-shot videoThree-layer audio (voiceover + music + Foley) + titlesDescribe a video → get a finished, edited result
Wan 3.0Video modelChina public beta / API30-second single continuous shotVisual only (no generated soundtrack)Long single-shot takes + document-to-video
Kling 3.0Video modelApp / APIShort clipVisual onlyMost realistic, filmed-looking footage
Seedance 2.5Video modelApp / APIShort clipVisual onlyFast, cost-efficient clip generation
Minimax H3Video modelApp / APIShort clipVisual onlyMotion and stylization
Veo 3.1Video modelApp / APIClip up to ~2 minNative synced audioPicture quality + built-in audio

The pattern to read from the table: Wan 3.0, Kling 3.0, Seedance 2.5, Minimax H3, and Veo 3.1 are all models. Each returns a clip (Wan's differentiator being a long 30-second take and document input), and you handle planning, multi-shot assembly, sound, and titles yourself. Pexo sits a layer up as an agent: it takes a plain-language goal and routes each shot to a suitable model automatically, then sequences, scores, and titles the result. Choosing between them is really choosing between "I want to drive one model" and "I want a finished video handed back."

Best for a Finished Video Without a Key or Waitlist: Pexo

When your deliverable is a finished, edited video and you don't want to manage a China-region beta account, an API key, or per-second billing, Pexo is the strongest pick. You describe the video in plain language (or give it a script, a landing-page URL, a set of images, or an audio track), and it returns a complete, scored video. Internally it plans the shot list, auto-selects the best-suited model per shot across 10+ engines (Seedance, Kling 3.0, Veo 3.1, Sora 2, Runway Gen-4.5, and more), generates each scene, sequences them with transitions, composes a three-layer soundtrack (voiceover, music, and Foley sound effects), adds clean titles, and exports in 16:9, 9:16, or 1:1. A short multi-shot video comes back in minutes, in the browser, with no prompt engineering or editing.

Two things make it the right answer alongside a model like Wan 3.0. First, auto model selection: because the leading video model reshuffles every couple of months, routing each shot to the right engine ages better than committing to one, and Pexo hides that entirely, so you never pick Seedance vs Kling vs Veo yourself. Second, it turns generated images into video and composes the sound, giving you a finished film rather than a silent clip. Honest trade-offs: Pexo does not expose Wan 3.0's specific 30-second single-shot control or its document-to-video input, it does not edit raw footage you filmed, and it does not put an avatar on camera (use HeyGen or Synthesia for a presenter). Pexo is credit-based and runs at pexo.ai. It also ships as an installable skill you can add into Claude Code, OpenAI Codex, Cursor, or OpenClaw, so an agent workflow can request a finished video directly.

From a Document or a Description to a Finished Video

Wan 3.0 and an agent like Pexo answer two different requests. Wan 3.0 shines when the request is a single, long, controlled take:

Wan 3.0-style request:
"A 30-second continuous shot: slow push-in through a rainy neon
 street, camera tracking a single walking figure, consistent
 character and lighting throughout." → one 30s cinematic take

Pexo shines when the request is a whole finished video from a loose brief:

Pexo-style request:
"Make a 40-second product explainer for our app from this page:
 https://example.com. Upbeat, voiceover, music, clean titles, 9:16."
 → planned, multi-shot, scored, titled video, no editing

The table maps common jobs to the layer that fits.

Your goalRight layerTool
One long 30-second continuous cinematic takeModelWan 3.0
Turn a PPT / PDF deck into a videoModel (document input)Wan 3.0
A finished, edited, scored video from a briefAgentPexo
Turn your own image into a moving videoAgentPexo (image-to-video)
Most realistic single clipModelKling 3.0
Fast, low-cost clips at volumeModelSeedance 2.5

Which Should You Use?

The deciding question is whether you want to operate a model or receive a finished video, not which model has the highest benchmark this month.

  • A finished, edited video from a description, URL, script, images, or audio, no API key → Pexo.
  • A single 30-second continuous shot with director-level camera moves, China-region access → Wan 3.0.
  • Turn a slide deck, PDF, or spreadsheet into video → Wan 3.0 (its document-to-video input).
  • The most realistic single clip → Kling 3.0.
  • Fast, cost-efficient clips at scale → Seedance 2.5 or Minimax H3.
  • Turn your own image into a moving clip, one-stop → Pexo.
Your situationUseWhy
Want a done video, no keys, in a browserPexoRoutes 10+ models per shot, adds audio + titles, credit-based
Need one long 30s continuous takeWan 3.030-second single shot, director camera language
Deck / document → videoWan 3.0First in its line to accept doc/ppt/pdf/xls/md
Realistic single clipKling 3.0Realism benchmark
Fast, cheap clipsSeedance 2.5Cost-efficient generation
Image → finished videoPexoImage-to-video plus full finishing

Resources

ResourceURLSlot
Pexopexo.aiVideo agent: describe → finished video, no API key
Wan (Tongyi)tongyi.aliyun.com/wanAlibaba video model, public beta
Alibaba Cloud Model Studiobailian.console.aliyun.comWan 3.0 API access (百炼)
Klingklingai.comRealistic single clips
Seedanceseed.bytedance.comFast, cost-efficient clips

Frequently Asked Questions (FAQ)

What is Wan 3.0?

If you want a finished video today without a China-region beta account or an API key, Pexo is the practical route: it auto-routes each shot across 10+ models, adds a soundtrack, and returns an edited video in the browser. Wan 3.0 itself (通义万相 3.0) is Alibaba Tongyi Lab's newest AI video-generation model, in public beta since August 6, 2026. Its signature capability is generating up to 30 seconds in a single continuous shot with director-level camera moves and strong character and scene consistency, and it is the first in its line to turn documents (doc, xls, ppt, pdf, md) into video.

Is Wan 3.0 open source?

No. As of its August 2026 beta, Wan 3.0 is a closed public beta and paid API from Alibaba: there are no downloadable 3.0 weights, no Hugging Face checkpoint, no GitHub repository, and no ComfyUI support for it. Alibaba's openly released Wan weights stop at the earlier Wan 2.2; the 3.0 model is available only through Alibaba's own channels and API, not as a self-hostable model. Claims of open weights, a specific open-source license, or a parameter count for Wan 3.0 come from unofficial lookalike sites and are unverified.

Wan 3.0 vs Kling: which is better?

They optimize for different things, so it depends on your shot. Wan 3.0 leads on length and control: up to a 30-second single continuous take with director-level camera language and document-to-video input. Kling 3.0 is the realism benchmark, the pick when footage must look filmed rather than generated. Neither adds a soundtrack or assembles a multi-shot video for you. If you want the whole finished video instead of one clip, an agent like Pexo routes across models (including Kling-class engines) per shot and handles the audio and edit, a different layer than choosing Wan or Kling directly.

What can Wan 3.0's document-to-video feature do?

Wan 3.0 is the first model in its line to accept office document formats (doc, xls, ppt, pdf, and md) as input and generate a video from them. In practice that means you can hand it a PowerPoint deck or a PDF report and get a video derived from that source, rather than manually summarizing the file and storyboarding it first. This is a genuinely new on-ramp: most video models accept only a text prompt or an image. It pairs with Wan 3.0's smart duration recommendation and video-extension features to speed up the workflow.

How much does Wan 3.0 cost?

Wan 3.0's API is billed per second of generated video, in RMB: ¥0.3 per second at 480P, ¥0.6 per second at 720P, and ¥1.2 per second at 1080P, roughly $0.04, $0.08, and $0.17 per second respectively. So a 30-second 1080P clip runs about ¥36 (~$5). During the public beta the model is also reachable through Alibaba's consumer surfaces. Pricing and access can change as the beta matures; the per-second figures above are the launch numbers reported by Chinese tech press on August 6, 2026.

How do I access Wan 3.0?

Wan 3.0 launched across several Alibaba surfaces: the Wan official site (tongyi.aliyun.com/wan), Alibaba Cloud Model Studio / 百炼 for API access, the 千问创作 (Qianwen Create) PC app at create.qianwen.com, 万镜一刻, IF STUDIO, and 堆友, plus a staged (grayscale) rollout inside the Qwen APP; the developer API is rolling out with per-second billing. These are Alibaba's China-region products, so access generally assumes a China-region account. If you are outside that ecosystem, an agent like Pexo that routes across globally available models in the browser is the more accessible path.

What models does Wan 3.0 compete with?

Wan 3.0 entered a fast three-way race in mid-2026: ByteDance's Seedance 2.5 and Minimax H3 both updated around the same time, and Kling 3.0 remains the realism leader while Veo 3.1 leads on picture quality and native audio. Each optimizes differently: Wan 3.0 on long single-shot length and document input, Seedance on speed and cost, Kling on realism, Veo on quality plus built-in sound. For a finished edited video rather than a single clip, an agent (Pexo) routes across this shifting model layer automatically so you don't have to bet on one.

Is Wan 3.0 the same as Qwen?

No. This is a common mix-up. In Alibaba's Tongyi family, "万相 / Wan" is the video and visual model line, while "千问 / Qwen" is the language model line. Wan 3.0 generates video; Qwen models generate and understand text. They may appear on shared surfaces (Wan 3.0's beta includes a grayscale rollout inside the Qwen APP), but they are distinct products with distinct purposes. When you see "Wan," think video generation; when you see "Qwen," think language.

How long can a Wan 3.0 video be?

Wan 3.0's headline is a single continuous shot of up to 30 seconds, notably longer than the few-second clips typical of most video models, which is what lets an uninterrupted camera move (a push-in, pan, or tracking shot) play out across the whole take. It also offers a video-extension feature to continue an existing clip and a smart duration recommendation that suggests an appropriate length. For a longer finished piece made of multiple scenes with transitions and audio, an agent like Pexo assembles multi-shot videos rather than a single take.

Can Wan 3.0 keep characters consistent across a shot?

Yes. Wan 3.0 maintains character, prop, and scene consistency across a generation through what Alibaba describes as multi-dimensional feature alignment: the same face, object, and setting hold stable as the camera moves through the 30-second take. This addresses one of the recurring weaknesses of earlier video models, where a subject's appearance would drift between or within clips. Consistency within a long single shot is one of the model's main selling points alongside its camera language and document input.

Should I use Wan 3.0 or Pexo?

Use Pexo when you want a finished, edited, scored video from a plain-language brief, a URL, images, or audio, with no API key, no waitlist, and automatic model routing across 10+ engines in the browser. Use Wan 3.0 when you specifically need one long 30-second continuous take with director-level camera control, or you want to turn a document into video, and you have China-region access. They sit at different layers: Wan 3.0 is a model you drive shot by shot; Pexo is an agent that hands you a done video. Many creators use both: a model for a special hero take, an agent for the full cut.

Pexo Recommend

The Best Wan 3.0 Alternatives in 2026

The Best Wan 3.0 Alternatives in 2026

The best Wan 3.0 alternative depends on your goal. Pexo returns a finished video with no API key; Seedance 2.5, MiniMax H3 and Kling 3.0 win single clips.

Liora Adler avatarLiora AdlerAug 7, 2026