FLUX 3 Ships 20-Second AI Videos With Audio Built Right In
All blog articles
Twenty seconds of AI video with real synced sound, no separate audio pass needed. That's the pitch behind FLUX 3, and it just turned Black Forest Labs from an image specialist into a genuine video contender overnight.
FLUX 3 turns an image specialist into a video rival
On July 23, Black Forest Labs unveiled FLUX 3, a multimodal foundation model trained jointly on images, video, and audio inside one architecture. The company describes it as a step toward visual intelligence models that perceive, predict, and act across environments.
This marks the first time the FLUX lineup, known until now for image generation, moves into motion. This release marks BFL's first public video generation model.
What FLUX 3 Video actually does
The model can create videos with native audio up to 20 seconds long in one generation, starting from text, images, keyframes, or reference clips. It also handles keyframe transitions, multilingual dialogue, and chains individual clips into longer multi-shot sequences.
CEO Robin Rombach summed up the logic behind the whole release in one blunt line: "the world is not made of still frames." Rather than bolting audio onto finished footage the way most rivals do, FLUX 3 trains sound and picture together from the start.
FLUX 3 vs Runway, Kling, and Grok Imagine
Black Forest Labs ran its own preference tests, and the self-reported numbers are bold. In preliminary internal evaluations, human reviewers preferred FLUX 3's output over Grok Imagine Video in 69% of comparisons, over Kling v3 Pro in 60%, over Runway Gen-4.5 in 77%, and over Luma Ray 3.2 in 93% of comparisons.
Grain of salt required here: it's the company grading its own homework, and independent testing hasn't happened yet. Still, if even part of that holds up, generators like Seedance 2.5 and Runway just got a serious new headache.
Robots on Audi's factory floor
The most concrete proof isn't a benchmark chart, it's a car plant. Audi's Christoph Schneider confirmed that FLUX-mimic is currently running in Audi production facilities, handling tasks including soft-body manipulation that conventional robotics cannot perform cost-effectively.
Turning the same backbone that renders a talking character into one that guides a robotic arm on an assembly line is an ambitious claim. Unlike the preference scores, though, this one is already running in production.
What's still missing from FLUX 3
FLUX 3 Image, arguably the feature most creators actually want, hasn't shipped. FLUX 3 Video is open now via gated early access; image generation follows in the coming weeks, and an open-weight FLUX 3 Dev backbone is planned for later.
Pricing, licensing, and a firm date for that open-weight Dev version remain unannounced. For now, most of us will just have to wait and see whether those flashy percentages survive contact with the real world.
What is FLUX 3 by Black Forest Labs?
FLUX 3 is a multimodal foundation model from Black Forest Labs that generates images and 20-second videos with native audio, and predicts robotic actions from one unified architecture.
Can I use FLUX 3 Video right now?
FLUX 3 Video is only available through a gated early access program as of July 2026, with API access, an image model, and an open-weight FLUX 3 Dev version planned for later this year.