Black Forest Labs unveils FLUX 3: images, video, audio, and robot actions from a single model
In one sentence Black Forest Labs launches FLUX 3, a unified multimodal model generating images, video up to 20 seconds with native synchronized audio, and — through the Action variant built with robotics startup mimic — action prediction for robots. Video and Action entered early access on July 23.
Until now, AI labs built one model for images, one for video, one for audio, perhaps another for robots. Black Forest Labs — the German lab founded by the creators of Stable Diffusion — took the opposite road: FLUX 3 is a single brain that learned all of these together, from one set of weights.
The flashiest piece is video: clips up to 20 seconds generated from text, images, or existing footage, with audio born inside the scene — dialogue, sound effects, and ambient noise, synchronized and multilingual. Clips can be chained while keeping characters consistent, transitions can be steered with keyframes, and multi-shot sequences can be assembled: editing building blocks, not just fragments.
The most unexpected part concerns robots. Together with Zurich-based startup mimic, the lab trained a variant called FLUX-mimic that predicts actions for robotic manipulation: the same model that imagines how the world moves in a video can suggest to a robot arm how to actually move. The robotics variant is reportedly already being tested on Audi production lines.
The underlying idea: a system that understands the physics of the world well enough to film it convincingly understands it well enough to act inside it. That is the world-model bet, and FLUX 3 is among the first to ship it as a product.
Companies
Black Forest Labs, mimic
Tools
FLUX 3
Tags
Sources