FLUX 3 renders 20-second clips with dialogue, sound effects and lip-sync built in
FLUX 3 is Black Forest Labs' multimodal foundation model: one architecture trained jointly on images, video and audio, generating all three from the same weights. A single clip runs up to 20 seconds with its soundtrack, including multilingual dialogue and lip-sync. Billing works pay-as-you-go, per second of generated video, with no subscription attached. The same backbone also predicts robot actions, currently tested on production tasks at Audi.
- 20-second clips rendered in one generation
- Native audio: synced dialogue, sound effects, ambience
- Lip-synced speech in several languages
- Draft mode at roughly a third of full price
- Pay-as-you-go billing, no seat fees
- Staged rollout: image model and open weights still pending
- Benchmarks are vendor-reported, awaiting independent testing
Twenty seconds of picture and sound in one pass
FLUX 3 Video generates clips up to 20 seconds long in HD or Full HD, with an optional soundtrack created alongside the frames: lip-synced dialogue, sound effects and background ambience. The model takes a text prompt, a starting image, a video to continue, or keyframes to steer controlled transitions.
In practice, you describe an aerial push over a rainy city and the file lands with camera motion, drizzle and traffic hum already mixed. The output is not locked to a cinematic look either: home camcorder footage, animation and animated typography all sit within reach. Against other text-to-video models, its strongest card is that native audio, delivered at no extra charge.
From video to Audi's robots: the FLUX 3 rollout
Black Forest Labs is releasing FLUX 3 one capability at a time, all built from the same foundation model. Video shipped first, initially in early access and then generally available through the API; the image model and an open-weight Dev release follow, each after its own testing phase.
The strangest part of the package is FLUX-mimic. This video-action model, co-developed with mimic robotics, drives robot manipulation in under 80 ms on a single RTX 5090, and Audi is trialing it on real factory tasks. Generated footage and robot arms run on the same brain, which is the whole thesis of the lab.
| Capability | What it produces | Where it stands |
|---|---|---|
| FLUX 3 Video | Clips up to 20 s with native audio | Available via the API and partners |
| FLUX 3 Image | Image generation and editing, multilingual text | Announced, rollout pending |
| FLUX 3 Dev | Downloadable open weights | Planned for a later stage |
| FLUX-mimic | Action prediction for robots | In testing with partners, including Audi |
What FLUX 3 costs and how to get started
FLUX 3 is billed on usage alone, with no subscription or per-seat licence: 1 credit equals $0.01 and each video is charged per generated second, based on length, resolution (HD or FHD) and render type. Draft mode, HD only, comes in at roughly a third of a full render, handy for exploring variants before paying full price for the keeper.
Pricing stays identical between the playground and the API (a small mercy: no markup sneaks in between your tests and production). For the per-second figures themselves, the rate card on bfl.ai is the reference, and it moves as the rollout progresses.
Frequently asked questions
Is FLUX 3 free?
No, FLUX 3 is billed on usage through a credit system, with no permanent free tier. There is no subscription: you pay for the seconds of video actually generated, and audio is included at no extra cost. An open-weight Dev release has been announced for a later stage, which will let you run the model on your own hardware.
Is FLUX 3 better than Kling or Seedance?
The only available numbers come from Black Forest Labs itself: in its human-preference tests on 10-second 720p clips, FLUX 3 won 93% of comparisons against Luma Ray 3.2, 77% against Runway Gen-4.5 and 60% against Kling v3 Pro, with a near coin flip against Seedance 2.0. Those results are preliminary and vendor-run, so treat them as a claim until third parties test it.
When will FLUX 3 Image and the open weights arrive?
The announced schedule places FLUX 3 Image after video, with the open-weight Dev release coming last. No firm dates have been published: every capability goes through an early-access phase with safety testing before general availability. Video generation, meanwhile, is already reachable through the API.
Does FLUX 3 support vertical video?
Yes, FLUX 3 sets its resolution bands by total pixels per frame rather than aspect ratio, so a vertical clip is treated like an equivalent 16:9 one. HD corresponds to about 1 megapixel per frame and FHD to 2 megapixels, while draft mode stays limited to HD.
Verdict: Twenty seconds of synced picture and sound per generation: that pitch speaks to ad studios, indie filmmakers and video teams building animatics, while developers can watch for the open-weight Dev release to bring the model onto their own machines.
