Black Forest Labs (BFL) has launched FLUX 3, a multimodal frontier model that moves the Freiburg, Germany-based lab well beyond its image-only roots. The new model generates images, video clips up to 20 seconds in length, and synchronized audio — all from a single prompt — with the same underlying architecture also targeting robotic vision and action prediction.
One Architecture, Four Products
Rather than stitching together separate image, video, and audio models behind a shared interface, BFL says FLUX 3 is jointly trained across modalities. The company frames this as a strategic bet: it wants enterprises to see creative generation, simulation, computer use, and robotics as expressions of a single capability it calls visual intelligence.
FLUX 3 ships — or will ship — across four product lines:
- FLUX 3 Video — with optional native audio generation; gated Early Access now
- FLUX 3 Image — rolling out in the coming weeks
- FLUX 3 Action — robotics and action prediction; gated Early Access now
- FLUX 3 Dev — open-weight, open source; arriving later this year
The gated early-access approach mirrors recent rollout strategies from Anthropic and OpenAI, though those delays were framed around safety and government coordination. BFL hasn't cited a specific reason beyond managing demand.
What's Missing at Launch
Several details enterprise buyers need simply aren't available yet:
- No published pricing for any FLUX 3 tier
- No production SLAs
- No downloadable weights — the open-weight FLUX 3 Dev arrives last in the sequence
- No image-model benchmarks — only preliminary video evaluation data
For developers who have come to expect a locally deployable FLUX variant alongside a major release, the delay is a meaningful step back. Open weights have been central to FLUX's adoption curve so far, and pushing them to the back of the queue changes the calculus for teams that depend on local inference.
Benchmark Numbers, With Asterisks
BFL published preference-test results from 10-second, 720p text-to-video comparisons, showing FLUX 3 was preferred over competitors at the following rates:
- Luma Ray 3.2 — 93%
- Runway Gen-4.5 — 77%
- Grok Imagine Video — 69%
- Kling v3 Pro — 60%
- Happy Horse v1 / v1.1 — 59% / 57%
- Seedance 2.0 — 52%
- Google Gemini Omni Flash — 52%
Critically, BFL itself labels these results a "preliminary evaluation of an early FLUX 3 candidate" — meaning the numbers describe a pre-release checkpoint, not the model entering early access. That caveat cuts both ways.
The strongest wins — against Luma Ray 3.2 and Runway Gen-4.5 — come against established products that aren't currently setting the pace in independent video rankings. The 52% result against Gemini Omni Flash is the number that matters most: Omni is the closest large-platform analogue to what FLUX 3 is attempting, and by BFL's own data, the two are statistically indistinguishable on short text-to-video quality.
Google's advantage in that matchup is concrete: Gemini Omni Flash is generally available via the Gemini API at $0.10 per second of generated 720p video — roughly $1.00 for a 10-second clip. FLUX 3 has no published price.
The Seedance 2.0 comparison, at 52%, is largely moot. ByteDance indefinitely postponed Seedance 2.0's international rollout after legal threats from Netflix, Warner Bros., Disney, Paramount, and Sony over alleged systematic copyright infringement — the model is currently inaccessible to most Western enterprises.
The Architecture Argument
BFL's technical foundation here is Self-Flow, its method for aligning multimodal understanding and generation within a single architecture, published in March 2026. The company says it scaled compute and training data substantially to cover video, images, and audio simultaneously, and that testing confirmed action prediction doesn't require a separate model foundation.
"You can't cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds." — Robin Rombach, co-founder and CEO, Black Forest Labs
Rombach's broader argument is that joint training creates compounding benefits: audio conveys timing and physical events that vision misses; language encodes goals that pixels can't easily express. Whether that theoretical coherence translates into measurable production advantages over modular approaches remains to be seen once full benchmarks and general availability arrive.
What This Means for Builders
For startup founders and product teams evaluating video generation infrastructure right now, the honest answer is: FLUX 3 isn't ready to evaluate yet. No pricing, no SLA, no open weights, and preliminary benchmarks against a pre-release checkpoint don't give teams enough signal to make procurement decisions.
The model's architectural ambition — extending a single foundation from media generation into robotic action prediction — is genuinely differentiated and worth tracking. But until general availability lands with published pricing and reproducible benchmarks, Gemini Omni Flash and Veo 3.1 remain the more actionable options for teams that need to ship now.



