Alibaba has announced the launch of its Wan2.2large video generation models.
In what the company said is a world first, the open-source models incorporate MoE (Mixture of Experts) architecture aiming to enable the production of cinematic-style video content with a single click.
The series features both the Wan2.2-T2V-A14B text-to-video and Wan2.2-I2V-A14B image-to-video models, as well as the hybrid Wan2.2-TI2V-5B providing support for either method within a single, unified framework.
Built on MoE architecture, Wan2.2-T2V-A14BMo and Wan2.2-I2V-A14B enable cinematic-grade creation, providing creators with precise control over dimensions including lighting, time of day, colour tone, camera angle, frame size, composition, focal length and more, said the company.
Both models enable the production of complex motions such as facial expressions, hand gestures and intricate sports movements, also delivering realistic representations with adherence to physical laws.
Aiming to address high computational consumption, Wan2.2-T2V-A14B and Wan2.2-I2V-A14B implement a two-expert design in the denoising process of diffusion models, including a high-noise expert focusing on overall scene layout and a low-noise expert to refine details and textures. Computational consumption is reduced by up to 50 per cent by activating only 14 billion of the available 27 billion parameters in each step.
The models were trained on what Alibaba said is a “substantially larger dataset” than Wan 2.1, with a 65.6 per cent increase in image data and 83.2 per cent increase in video data.
Utilising a high-compression 3D VAE architecture to achieve a temporal and spatial compression ratio of 4x16x16, the hybrid Wan2.2-TI2V-5B model enhances the overall information compression rate to 64 and can generate a 5 second 720P video in minutes, using a consumer-grade GPU.
The models are available on Hugging Face and GitHub, as well as Alibaba Cloud’s open-source community, ModelScope.