The rapid pace of progress in large language models (LLMs) has spilled into nearly every corner of generative AI.
What started as text systems capable of coherent conversation quickly, popularized since the release of ChatGPT from OpenAI, has expanded into image and then video generation, with new capabilities appearing in months rather than years.
Western labs set much of the early direction, yet Chinese teams have closed the gap and in several practical dimensions now match or exceed it.
Alibaba's Wan series illustrates the shift.
Earlier versions already delivered competitive quality and open weights for some releases.
'Wan 3.0,' now in public beta, extends that trajectory with longer native clips and broader input handling.
Wan 3.0 generates video up to 30 seconds in a single pass, doubling the previous Wan 2.7 limit of 15 seconds. Resolutions range from 480p to 1080p with adaptive aspect ratios.
The model accepts text, images, audio and video as inputs and adds structured documents such as PDFs, PowerPoint files, spreadsheets, Markdown and web pages.
It can read these materials, extract relevant content and turn them into coherent video sequences.
Character consistency is described as production-grade, with more expressive faces, stable identity across frames and improved rendering of digital interfaces and motion graphics.
Audio is generated together with the visuals, producing synchronized dialogue, sound effects and ambient tracks in the same output file.
Reference control allows users to lock appearance, style or spatial layout from multiple sources in one request.
By comparison, OpenAI’s Sora 2 reaches 20 seconds at 1080p on its higher tier and includes native audio, yet the consumer application has already shut down and the API is scheduled for full deprecation later in 2026.
Google’s Veo 3.1 produces up to 8-second clips in a single generation that can be extended, supports native audio and offers 4K output, but the base length remains shorter.
Runway’s Gen-4.5 focuses on fine-grained control tools such as motion brushes and reaches roughly 10 seconds in core generation modes, with limited native audio.
The practical difference for Wan 3.0 shows most clearly in workflows that begin with existing materials rather than pure prompts or image sets.
A product brief in a slide deck or a technical document can be fed directly into the model to produce a 30-second explainer without intermediate conversion steps.
Character and product fidelity remain stable across the longer duration, reducing the need for post-generation stitching that shorter models such as Alibaba's own HappyHorse 1.0 often require.
ByteDance's Seedance 2.0 matches the 30-second length and exceeds it in raw reference count, yet lacks the document and webpage ingestion that defines Wan 3.0’s omni-reference approach.
But when compared to Kuaishou's Kling 3.0’s native 4K and multi-shot tools provide higher spatial resolution and more explicit shot planning within its 15-second window
The practical difference shows most clearly in workflows that begin with existing materials rather than pure prompts.
A product brief in a slide deck or a technical document can be fed directly into Wan 3.0 to produce a 30-second explainer without intermediate conversion steps.
Character and product fidelity remain stable across the longer duration, reducing the need for post-generation stitching that shorter models often require. Pricing on the public beta sits at 0.05 dollars per second for 480p, 0.10 for 720p and 0.20 for 1080p, with access available through Alibaba Cloud Model Studio and Qwen Cloud while full API rollout continues.
Independent leaderboard data for Wan 3.0 itself is still limited because the model only recently entered beta.
Earlier Wan releases ranked competitively on public arenas but did not dominate the absolute top positions.
The advance to 30-second native generation and document-level omni-reference nonetheless moves the series into a distinct operating range that few Western systems currently occupy.
The overall pattern is consistent with the wider industry: rapid iteration continues on all sides, and the functional lead in any single capability can shift within a single model cycle.
The significance of Wan 3.0 is therefore less about declaring a single winner than about how quickly the boundaries between competing video models are moving. A
year ago, capabilities such as longer coherent generation, native audio and multimodal reference control were differentiators; increasingly, they are becoming baseline expectations.
For creators, the more important shift may be the move from video generation to video production. Models that can understand documents, images, audio, webpages and existing footage are beginning to function less like prompt-driven clip generators and more like production tools that can transform source material into finished sequences.
Wan 3.0 is an important example of that transition, given by Alibaba's notable strong position in an increasingly crowded field.
Now, that many major labs are into the same game, the lead is becoming increasingly fluid, and the next major advantage may belong not to the model that generates the prettiest clip, but to the one that can turn the most information into usable video with the fewest steps in between.




















































































































































































































































































































































































