Alibaba has launched the public beta version of Wan3.0, its latest AI video generation model capable of producing clips up to 30 seconds long. The new model supports comprehensive multimodal reference inputs, allowing users to convert static, text-heavy data directly into dynamic video content for professional workflows.
Key Feature Highlights of Alibaba Wan3.0
Compared to mainstream AI video generators that produce clips lasting only a few seconds to 15 seconds—such as its predecessor, Wan2.7-Video—Wan3.0 doubles output length to 30 seconds per clip. This native extension enables complex camera movements, continuous unbroken shots, and longer narrative timelines. It also incorporates an intelligent duration feature that recommends optimal video lengths based on user prompts.
Wan3.0 supports simultaneous multimodal inputs, including text, image, video, audio, web pages, and documents such as PDFs and PowerPoint presentations. The model addresses visual drifting and distortion through high-precision visual continuity, rendering realistic human faces with synchronized micro-expressions, natural multilingual voice outputs, and stable motion graphics. It strictly maintains control over character features, product details, spatial layouts, and audio consistency across scenes.
Access and Availability
Applications for Wan3.0 model testing are currently open for users during its public beta phase. Testing access is available via Alibaba Cloud’s AI development platform Model Studio and the AI-native cloud platform Qwen Cloud.
#Alibaba #Wan3 #GenerativeAI #AIVideo #AlibabaCloud #ModelStudio #QwenCloud

