Google releases Veo 3.2: 4K video generation at 60fps with native lip-sync and integrated audio
In one sentence Google DeepMind launches Veo 3.2, advancing AI video generation to 4K 60fps with native lip-sync, multi-character scene consistency, and integrated audio generation, available via Gemini API and Vertex AI.
Imagine describing a scene in plain words and getting back a high-quality video in minutes, complete with people speaking naturally, background music, and ambient sounds, all generated automatically by an AI system. That is what Google Veo 3.2 promises.
The previous version of Veo was already impressive, but it had some clear limitations: videos were often low resolution, characters would change appearance between scenes, and lip movements did not match the spoken audio. Veo 3.2 addresses all of these issues at once.
With the new model, you can generate videos up to 4K at 60 frames per second, a quality comparable to professional film production. The system understands who the characters are and keeps them consistent throughout: if a red-haired woman appears in the first scene, she will look exactly the same in every subsequent scene. Native lip-sync means that when a character speaks, the lip movements match the audio precisely, just like in a real film.
On top of that, Veo 3.2 automatically generates audio as well: dialogue, sound effects, and ambient music are all created alongside the video, with no need for separate tools.
Who can use it? Developers and businesses can access the model through the Gemini API or Google's Vertex AI cloud platform. The product positions itself as a direct competitor to OpenAI's Sora 2, raising the bar further in the race for AI video generation.
Companies
Google DeepMind
Tools
Veo 3.2, Vertex AI
Tags
Sources