Black Forest Labs launches FLUX 3—a multimodal model that generates video with its own audio
FLUX 3 jointly learns images, video and audio, and can produce 20-second clips with native sound—though only in limited release for now.
Black Forest Labs launches FLUX 3—a multimodal model that generates video with its own audio
FLUX 3 jointly learns images, video and audio, and can produce 20-second clips with native sound—though only in limited release for now.
In Brief
- FLUX 3 is Black Forest Labs’ new multimodal foundation model, trained across images, video and audio.
- It generates 20-second video clips with native audio, the company says, though the full vision ships in limited early access.
- The release pressures rivals in the image-to-video race as 2026 compresses the field.
Black Forest Labs unveiled FLUX 3 on July 23, its first model to treat images, video and audio as one training surface rather than separate products. The German lab behind the original FLUX image models says the system can generate short video clips—up to 20 seconds—with sound synthesized natively rather than dubbed on afterward.
The move extends a year of rapid escalation in generative video, where labs have raced from still images to minutes of coherent motion. IBTimes notes the “product shipping today is much narrower than the announcement suggested,” a familiar gap between demo and delivery.
For a company that built its reputation on still-image quality, FLUX 3 is a declaration that the next battleground is unified audiovisual generation.
What FLUX 3 actually does
Per Black Forest Labs, FLUX 3 “jointly learns from images, video, audio” and is positioned as the “backbone of visual intelligence.” The model is the lab’s first to span all three modalities in a single architecture.
GlobeNewswire carries the company’s launch statement, describing FLUX 3 as “a new multimodal frontier model” for visual intelligence. The release emphasizes training across combined datasets rather than stitching separate models together.
IBTimes adds that the system was unveiled “as a unified AI architecture trained across images, video, audio and physical action,” though it cautions that not every promised capability is in the box at launch.
Where it fits in the race
The model enters a crowded field of image-to-video systems, where the differentiator is no longer “can it make a clip” but “can it keep sound, motion and intent consistent.” Native audio is the headline differentiator Black Forest Labs is leaning on.
Black Forest Labs is releasing FLUX 3 in early access rather than general availability, a deliberate gate while it watches for misuse and capacity strain. Most third-party developers will wait weeks for broad API access.
The practical question for customers is access and cost. A narrower day-one product means benchmarks and creator workflows will determine whether FLUX 3’s unified promise survives contact with production use.
FAQ
What can FLUX 3 generate?
FLUX 3 produces images and up to 20-second video clips with native audio, trained on a combined image, video and audio dataset.
Is FLUX 3 publicly available?
No—Black Forest Labs released it in limited early access, not general availability.
How does it compare to rivals?
Black Forest Labs positions FLUX 3 as a unified multimodal model, but independent testing of the shipped product is still limited at launch.