Black Forest Labs’ Flux 3 generates 20-second videos with native audio — and beats rivals in early tests
Black Forest Labs' Flux 3 generates 20-second videos with native audio and beats rivals in early tests — plus a robotics action model tested at Audi.
In Brief
- Black Forest Labs released Flux 3, a multimodal foundation model trained on images, video, and audio that generates clips up to 20 seconds long with native sound.
- In early evaluations, Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons and over Runway Gen-4.5 in 77 percent, though margins narrow against Kling v3 Pro and Seedance 2.0.
- The company also developed Flux-mimic, a video action model for robotics applications that is already being tested at Audi.
German AI company Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio simultaneously and generates videos up to 20 seconds long with native audio — a first for the Stability AI spinout best known until now for its Flux image models. According to The Decoder, the model beat several rivals in the company’s early video generation tests.
The system supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips that chain individual shots into longer multi-shot sequences. Black Forest Labs says the model is especially strong at rendering human facial expressions and matching sounds to the physical events that produce them.
The release positions the company in the broader race toward world models — systems that understand physical reality rather than merely rendering it, a thesis Google DeepMind researchers have also advanced. BFL describes Flux 3 as a step toward “real-world visual intelligence,” which it defines as models that can “perceive, predict, and act across physical and digital environments,” per The Decoder.
What Flux 3 Can Do — and How It Stacks Up
In early evaluations using 10-second clips at 720p, Black Forest Labs reports that Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. The margins narrow against stronger competitors: Flux 3 won 60 percent of matchups against Kling v3 Pro and 52 percent against both Seedance 2.0 and Gemini Omni Flash.
The company acknowledges the results are preliminary, and no independent tests are available yet. A 50 percent score in these head-to-head preference comparisons indicates a tie, so the single-digit edge over the strongest rivals suggests Flux 3 has reached the leading pack rather than lapped it.
Beyond video, BFL expects Flux 3 to improve image generation, particularly for complex prompts and accurate text rendering in multiple languages. Flux 3 Image is set to enter early access within the next few weeks, and the company plans to release open-weight access to the multimodal backbone under the name “Flux 3 Dev” — a notable commitment as rivals increasingly lock down frontier creative AI tooling.
From Video Generation to Robot Hands
The architectural bet behind the release is that no single modality captures reality in full. Images show spatial structure, video captures how it changes over time, and audio can reveal links between mechanical events and the sounds they produce. Training on all three together lets each modality fill gaps for the others — an approach BFL says delivers better results than the previously standard flow-matching method, both in generation quality and in the model’s grasp of the physical world.
That physical grounding feeds directly into robotics. A dedicated action component in the multimodal transformer provides the foundation for embodied applications, and the company worked with Mimic Robotics to develop Flux-mimic, a video action model that is already being tested at carmaker Audi, according to The Decoder.
Black Forest Labs is rolling out capabilities in stages, with early-access phases for feedback and safety testing, while action prediction will initially be offered through select partners. Longer term, the company says it is working on next-generation models that combine perception, action, and language prediction in a single system.
FAQ
How long are the videos Flux 3 can generate?
Up to 20 seconds with native audio — a first for Black Forest Labs — with support for text-to-video, image-to-video, keyframe transitions, and multi-shot sequences chained by agents.
How does Flux 3 compare with rival video models?
In BFL’s early tests on 10-second 720p clips, it was preferred over Luma Ray 3.2 in 93 percent of comparisons and over Runway Gen-4.5 in 77 percent, but the company says results are preliminary and no independent benchmarks exist yet.
Will Flux 3 be available as open weights?
Black Forest Labs plans open-weight access to the multimodal backbone under the name “Flux 3 Dev,” while Flux 3 Image enters early access in the coming weeks.