Alibaba Opens Wan3.0 Beta: 30-Second AI Video Built From Text, Images — Even PDFs

alibaba opens wan3 0 beta 30 second ai video built from text images even pdfs The figure to focus on with Alibaba's latest video model is thirty seconds — precisely twice the runtime Wan2.5 was capable of producing.

The figure to focus on with Alibaba’s latest video model is thirty seconds — precisely twice the runtime Wan2.5 was capable of producing.

Now available in beta, Wan3.0 gets stranger once you look at what it accepts on the way in. Text is table stakes. PDFs, web pages and PowerPoint files are not, yet the model digests them and spits out moving footage.

Ten images, five videos, five audio clips, one prompt

Text, images, video and audio can all be fed to Wan3.0 simultaneously. One prompt has room for as many as 10 images, five videos and five audio clips.

Juggling that volume of reference material without losing the thread is, in fairness, the central difficulty of generative video. Faces drift. Interfaces dissolve into gibberish. Props morph from one shot to the next.

According to Alibaba, Wan3.0 is better at preserving detail pulled from reference material — characters, props and spatial layouts in particular. That is the company’s claim. Visual drift remains the failure mode that every model in this space is still wrestling with, and a vendor asserting improved consistency is not proof that the consistency survives a full 30-second clip.

It picks the length for you

Instead of leaving you to guess, the model proposes a video length derived from your prompt. An extension tool is also included for pushing existing videos out to a longer runtime.

Neither is a headline feature. But anyone who has watched a generated clip cut off mid-gesture understands exactly why they exist.

Two tiers, one discount

Access comes via the wan.video site, through Alibaba Cloud Model Studio, or over the API on Qwen Cloud. Two tiers are on offer: Standard, discounted by 30 percent at the moment, plus a quicker Prime option.

The target market Alibaba describes is broad. Faster film production. Short-form dramas. Social clips made by creators. Companies converting text and images into marketing and training video. And developers producing realistic simulation footage to train autonomous vehicles and robotics systems — the sole use case on the list where nobody has to actually enjoy watching the result.

The money behind it

Some context for the timing: Alibaba’s AI spending is running hot. To bankroll that effort, the company recently announced the biggest share sale ever undertaken by a Hong Kong-listed company.

Last week, meanwhile, it posted a 75 percent year-over-year decline in quarterly profit, a drop attributed to steeply increased AI investments.

Wan3.0, in other words, comes with an invoice attached. Anyone trialling it should begin on the Standard tier while the 30 percent discount lasts — and feed it a document before feeding it a text prompt. Document-to-video is the capability none of the rival models on your shortlist are promoting.