Last updated: July 24, 2026
Black Forest Labs (BFL) launched FLUX 3 on July 23, 2026, and the headline is not “a better image model.” The headline is that the company behind FLUX.1 and FLUX.2 has rebuilt its flagship as a multimodal foundation model — one architecture trained jointly on images, video, and audio, and even extended to robot actions. FLUX 3 is in Early Access now, with FLUX 3 Video available first.
Quick answer: FLUX 3 is BFL’s first model built on the idea that a single network should learn a representation of the world from every modality at once, not just pixels. It generates video with native, synchronized audio up to 20 seconds in one generation, synthesizes and edits images, and — through a collaboration with mimic robotics — powers FLUX-mimic, a video-action model already being tested and deployed on Audi production lines. BFL’s early, vendor-run comparisons put FLUX 3 Video ahead of several established video models on human preference, but those results are explicitly preliminary. There is no public price yet: access during Early Access is by request.
| FLUX 3 fact | Detail |
|---|---|
| Release date | July 23, 2026, per BFL’s announcement |
| Status | Early Access; FLUX 3 Video live first, model page still labeled “Coming Soon” for general access |
| What changed | FLUX 1/2 generate images; FLUX 3 is one multimodal model for image, video, audio, and action |
| Video | Up to 20 seconds with native audio in a single generation; text-, image-, video-, and keyframe-conditioned |
| Image | Synthesis and editing; Early Access “in the following weeks,” per BFL |
| Action / robotics | FLUX-mimic, a video-action model with mimic robotics, deployed at Audi |
| Foundation | Built on BFL’s Self-Flow research; scaled-up compute and data |
| Open weights | FLUX 3 Dev (multimodal backbone) announced, not yet released |
| Pricing | Not publicly listed; Early Access is by request |
| Evidence level | Vendor-run preliminary preference evaluations; no independent benchmarks yet |
Source check — July 24, 2026: this article checks BFL’s official FLUX 3 announcement, the companion FLUX 3 x mimic post, the FLUX 3 model page, BFL’s pricing page (which still lists only the FLUX.2 lineup), the Self-Flow research page, and mimic’s FLUX-mimic post. Because FLUX 3 is in Early Access and BFL says evaluations are still improving, treat every capability and benchmark claim as provisional and verify the live model page before planning production work.
For where FLUX.2 fits among today’s image generators, see Top AI Image Generators in 2026. For the Google video model FLUX 3 is directly compared against, see Gemini Omni Flash, and for Google’s image side, Nano Banana 2 Lite.
What FLUX 3 Is — and Why It Is Not Just “FLUX 2 but Better”
The cleanest way to understand FLUX 3 is the contrast BFL itself draws: FLUX 1 and FLUX 2 generate images. FLUX 3 expands into multimodality.
FLUX 3 is built on Self-Flow, BFL’s approach for aligning multimodal generation and understanding inside one architecture. BFL says it then “significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.”
The argument in BFL’s announcement is worth quoting because it explains the product shape: no single modality fully describes the world. Images capture spatial structure at one moment; video restores time and physical dynamics; audio reveals causal links between mechanical events and sound; language ties all of it to goals and instructions. Train on all of them together, BFL argues, and their mutual constraints become evidence about one underlying reality — “the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past.”
That is a different bet from the FLUX.2 family, which is an image lineup. For the standing FLUX.2 picture — the [pro], [max], [flex], [dev], and [klein] tiers — see the FLUX.2 section in Top AI Image Generators in 2026.
FLUX 3 Video: 20 Seconds With Native Audio
Video is the lead capability, and BFL is candid that it is also the hard part — video prediction accounts for over 95% of total compute costs in the FLUX 3 training run, while audio is less than 0.5% of the tokens in a 720p video with audio.
FLUX 3 Video generates clips up to 20 seconds with native audio in a single generation. BFL lists the core capabilities as:
- Text-to-video generation;
- Image-to-video, either animating from a starting frame or using images as visual references;
- Video-to-video from a reference clip, carrying central elements (for example the same character) into a new scene;
- Generative video-audio continuation from input video and audio;
- Keyframe-to-video for controlled transitions between defined moments;
- Multilingual dialogue;
- A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics;
- Agentic chaining of individual clips into longer, multi-shot sequences;
- Strong typography generation and animated designs.
BFL highlights that FLUX 3 Video is “already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities,” and that clips can be combined into sequences lasting several minutes while visual references keep characters consistent across scenes.
The early benchmark picture (read it as vendor preference data)
For its preliminary analysis, BFL generated 10-second text-to-video clips in 720p with audio and ran human-preference comparisons. BFL reports FLUX 3 was preferred:
| Compared against | FLUX 3 preferred in |
|---|---|
| Luma Ray 3.2 | 93% of comparisons |
| Runway Gen-4.5 | 77% |
| Grok Imagine Video | up to 69% |
| Kling v3 Pro | 60% |
| Happy Horse v1 | 59% |
| Happy Horse 1.1 | 57% |
| Seedance 2.0 | 52% |
| Gemini Omni Flash | 52% |
Two honest caveats, both from BFL’s own framing:
- “Evaluations are early and we expect further improvements.” BFL says the model and the harness around it are still in development, and these results are preliminary.
- These are vendor-run preference comparisons, not an independent benchmark suite. A 52% preference over Gemini Omni Flash is essentially a coin-flip-plus-margin, not a clear win, and the lopsided numbers against Luma and Runway should be re-checked once independent evaluators run the same models.
FLUX 3 Image: Coming in the Following Weeks
FLUX 3 also synthesizes and edits images across styles, aspect ratios, and resolutions. BFL says that during midtraining, FLUX 3 already showed “a significant improvement over earlier versions of FLUX,” particularly in handling complex prompts and rendering high-accuracy text in multiple languages.
The important scheduling detail: FLUX 3 Image Early Access opens “in the following weeks,” not at launch. So at the time of writing, the image capability is demonstrated but not yet broadly accessible, while FLUX 3 Video is the live Early Access product.
FLUX-mimic: The Same Backbone, Now Driving Robots at Audi
The most surprising part of the FLUX 3 story is not video — it is that an early version of FLUX 3 is already running on robots.
Together with mimic robotics, BFL built FLUX-mimic, a video-action model that combines the FLUX 3 backbone with mimic’s robot-learning and deployment stack. Per BFL’s companion post, FLUX-mimic has been tested and deployed at Audi, handling factory tasks that conventional automation struggles with: kitting parts into structured trays, inserting electronic control units into tight fixtures, assembling components, and handling soft, flexible materials like seals and cables.
The technical claim is that action prediction is “one more view of the reality” the video model already learns. BFL reports that when action prediction was added to the training curriculum, human ratings on video generation initially fell by up to 10%, then recovered to full quality after about 3,500 steps while the model also gained action prediction — i.e., adding the action modality did not permanently cost video capacity.
Other FLUX-mimic details from BFL and mimic:
- The action decoder outperforms previous vision-language-action (VLA) models even with a completely frozen FLUX backbone, and reaches state-of-the-art success rates when the backbone is finetuned alongside the decoder;
- mimic reports up to 10x sample efficiency for video-action models over VLA models, and FLUX-mimic combines that with the Self-Flow representation gains;
- The backbone runs input-to-world-representation in under 80ms on a single NVIDIA RTX 5090, with the full robot system reaching 101ms reaction times — on the order of human visual reaction time.
Audi’s Christoph Schneider (Production Lab) is quoted saying the robots “solve complex soft-body manipulation work that would have been simply impossible with conventional robotics.”
This matters beyond robotics because it is BFL’s proof of thesis: content creation and physical AI are two applications of the same foundation model. If a model must learn how the world behaves to generate convincing video, that same world model can be decoded into robot actions.
The FLUX 3 Launch Plan: Four Capabilities, Staged
BFL is rolling FLUX 3 out in stages, each after an Early Access phase with safety testing. All four capabilities come from the same underlying multimodal flow-matching model:
| Capability | What it is | Status |
|---|---|---|
| FLUX 3 Video | Video and audio generation and editing via API and private weight access | Early Access live |
| FLUX 3 Image | Image synthesis and editing via API and private weight access | Early Access “in the following weeks” |
| FLUX 3 Action / FLUX-mimic | Action prediction through selected research and commercial partners, starting with mimic robotics | Partner Early Access |
| FLUX 3 Dev | Open-weight access to the multimodal backbone for content creation and action prediction | Announced, not released |
BFL also says it will release more technical details on the underlying approach. For teams that want open weights, FLUX 3 Dev is the one to watch — but it is promised, not available, at the time of writing.
Pricing: Not Public Yet
Unlike the FLUX.2 lineup, which BFL sells pay-as-you-go through its API and licenses for self-hosting (see the pricing page and the FLUX.2 tiers in our image tools guide), FLUX 3 has no published price. The pricing page’s calculator still covers only FLUX.2 [flex], [pro], [max], and [klein], plus FLUX Tools.
During Early Access, the path in is BFL’s request form. If history is a guide, FLUX 3 will likely land in the same API-plus-licensing structure as FLUX.2 once it reaches general availability, with FLUX 3 Dev carrying the open-weight role that FLUX.2 [dev] and [klein] hold today — but that is an inference from the existing lineup, not a published FLUX 3 price.
FLUX 3 vs the FLUX.2 Family
For anyone already building on FLUX.2, the practical question is what FLUX 3 changes.
| Dimension | FLUX.2 family | FLUX 3 |
|---|---|---|
| Core modality | Image generation and editing | Image, video, audio, and action in one model |
| Video | Not the focus of the FLUX.2 image lineup | Native video with synchronized audio, up to 20s |
| Audio | No | Native, generated with the video |
| Action / robotics | No | FLUX-mimic and FLUX 3 Action via partners |
| Open weights | [dev] (32B) and [klein] (4B/9B) available | FLUX 3 Dev announced, not yet released |
| Availability | GA via API and licensing | Early Access by request |
| Pricing | Public, cents-per-image range | Not published yet |
The short version: FLUX.2 remains the production-ready, priced, open-weight image option today. FLUX 3 is the multimodal successor, available first for video and only through Early Access. They are not a clean swap yet.
Who Should Request Early Access
- You produce short-form video and want native synchronized audio
- You already build on FLUX.2 and want the upgrade path early
- You work in physical AI / robotics and want a video-action backbone
- You need stable GA pricing before committing
- Your pipeline depends on open weights (wait for FLUX 3 Dev)
- You need independent benchmarks, not vendor preference data
- You need production SLAs and procurement-ready terms now
- You only need still images and FLUX.2 already covers you
- You cannot operate on a model that BFL says is still improving
Final Verdict: The Interesting Story Is the Unified Model, Not the Benchmark Table
FLUX 3 is best read as a thesis, backed by an Early Access product. Black Forest Labs is arguing that the same foundation that generates convincing video must already understand how the world behaves — and that this single world model can drive image and audio generation and robot actions. FLUX-mimic running on Audi production lines is the strongest evidence for that claim, far more than the preliminary video preference table.
For creators, the practical takeaway is narrower: FLUX 3 Video with native audio up to 20 seconds is worth evaluating if short-form, sound-synchronized video is your job, and the preference numbers — especially against Luma Ray 3.2 and Runway Gen-4.5 — are promising enough to justify an Early Access test. But the numbers are vendor-run and provisional, pricing is not public, and the open-weight FLUX 3 Dev is still ahead.
The right posture: watch FLUX 3 as the direction of travel for multimodal generation, request Early Access if video-with-audio or physical AI is your domain, and keep FLUX.2 as the priced, open-weight production option until FLUX 3 reaches general availability.