What Is FLUX 3? Multimodal Model Explained

Guide · 11 min read · Published: · Updated:

What is FLUX 3? Black Forest Labs' first multimodal model makes 20-second video with native audio. Here's what is real, what is hype, and how to get access.

MoClaw Editorial · MoClaw editorial team
What Is FLUX 3? Multimodal Model Explained
Table of Contents

Share this

FLUX 3 is Black Forest Labs' first multimodal model: a single system that generates images, video up to 20 seconds, and synchronized native audio, announced on July 23, 2026 (Black Forest Labs). It is not an image-only tool, and it is not unreleased, two claims still circulating from a placeholder page that leaked days before launch.

The most revealing number in the announcement is a training one. BFL says video prediction accounts for over 95% of FLUX 3's total training compute, while audio makes up less than 0.5% of the tokens in a 720p video with audio (BFL x mimic). The model learned to see motion the hard way and picked up sound almost as a side effect. That single design choice explains most of what FLUX 3 is, and most of what it is not.

Key Takeaways:

  • FLUX 3 is a unified multimodal model, image plus video plus audio in one architecture, not three separate models behind one badge. It was announced July 23, 2026.
  • Only FLUX 3 Video is in early access today, and it is approval-gated with no public API yet.
  • Native synchronized audio in clips up to 20 seconds is the headline feature, a first for Black Forest Labs.
  • Benchmark win rates are BFL's own numbers measured on a pre-release candidate, so read them as a floor, not a verdict.
  • If you saw that FLUX 3 is "unreleased" or "image-only," that is stale leak-era information from a placeholder page, and it is wrong.

What FLUX 3 Actually Is: One Model, Three Modalities

The cleanest way to understand FLUX 3 is by what it refuses to be. Most "multimodal" launches are a bundle: an image model, a separate video model, and a bolted-on audio pass, sold under one name. BFL's framing for FLUX 3 is the opposite. The model "jointly learns from images, videos, and audio within a unified architecture," and the company positions the whole effort as a step "towards multimodal flow models as the backbone of visual intelligence" (Black Forest Labs).

The training method has a name, Self-Flow, which BFL describes as its approach for aligning multimodal generation and understanding inside the same underlying model. You do not need the math to grasp the consequence. Because image, video, and audio share one backbone, the same system that renders a still frame also understands how that frame should move and what it should sound like. That is the bet: that generation quality across modalities compounds when they are learned together rather than stitched together.

Whether that bet pays off in your hands is a separate question, and today most people cannot test it, which we get to below.

What this resolves: FLUX 3 is a single multimodal model, not a marketing umbrella over three.

What it leaves unsolved: "Unified architecture" is an engineering claim, not a quality guarantee. The proof is in outputs the public mostly cannot generate yet.


What Is Actually New in FLUX 3

Strip the launch copy down to the parts that change what you can make, and four things stand out.

Twenty-second video with native audio. FLUX 3 produces clips up to 20 seconds long, and "all outputs come with native audio generation," meaning dialogue, sound effects, and ambient sound generated alongside the picture and synced to it (Black Forest Labs). Native synchronized audio is a first for BFL, whose earlier FLUX models were image-only.

Clip chaining for longer sequences. A single clip caps at 20 seconds, but FLUX 3 supports "agent-driven links between clips for longer multi-shot sequences," alongside text-to-video, image-to-video, video-to-video, and keyframe transitions (The Decoder). The word "agent-driven" is not decoration, and it matters for where this technology is heading.

Sharper image editing. The image side inherits the FLUX lineage that already powers editing inside tools like Adobe Photoshop and Picsart, now folded into the unified model rather than shipped as a standalone checkpoint (Yahoo Finance).

Action prediction for robots. A research branch called FLUX-mimic extends the same video backbone into a video-action model. In partnership with the robotics firm mimic, BFL says FLUX-mimic has been "testing and deploying" on real factory use cases, with Audi named as a deployment site (BFL x mimic). This tier is not something a content team will touch, but it tells you the model is aimed at physical-world prediction, not just pretty clips.

What this resolves: The genuinely new capabilities are native audio, clip chaining, and an action model, not just "better images."

What it leaves unsolved: Every one of these is demonstrated in BFL's own materials. Independent, hands-on verification is thin because access is narrow.

What builders said about FLUX 3 in the first days after launch (real posts from X):
  • @bfl_ai profile photoBlack Forest Labs@bfl_ai

    Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. FLUX 3 Video is now available in early access (link below). Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.

    ❤ 6,074 · 👁 1.1M

  • @dreamingtulpa profile photoDreaming Tulpa 🥓👑@dreamingtulpa

    seriously impressed by how flux 3 can do dialogue with super vague prompts like "ranting about ai"

    ❤ 2,731 · 👁 245K

  • @venturetwins profile photoJustine Moore@venturetwins

    I've been playing around with FLUX 3 for the last few weeks - it's an incredibly impressive model. It can generate up to 20 seconds of video across multiple shots with one prompt 🤯

    ❤ 638 · 👁 59K

  • @AIWarper profile photoA.I.Warper@AIWarper

    FLUX 3 for me feels like a better SORA. I think people are going to make some really wacky stuff with this model. “Girl from the attached ref images is break dancing on a flattened cardboard box on a bustling NYC sidewalk in 1993”

    ❤ 839 · 👁 43K

  • @jerrod_lew profile photoJerrod Lew@jerrod_lew

    FLUX 3 is here. Been testing it today, and it is outstanding. 20 second generations, does very well with detailed prompts and with amazing cinematography.

    ❤ 602 · 👁 42K


Why Video Was the Hard Part, and Audio Was Almost Free

This is the section other launch coverage skipped, and it is the most useful thing in the announcement for understanding the model's shape.

BFL is unusually candid about where the effort went. Video prediction accounts for "over 95% of the total compute costs," while audio comprises "less than 0.5% of the tokens in a 720p video with audio" (BFL x mimic). Read those two numbers together and a design philosophy falls out: the company spent almost everything teaching the model to predict how the world moves, and treated sound as a nearly free consequence of that.

Where FLUX 3's training compute went: over 95% to video prediction, under 0.5% of tokens to audio. Source: Black Forest Labs.
Where FLUX 3's training compute went: over 95% to video prediction, under 0.5% of tokens to audio. Source: Black Forest Labs.

BFL states the logic directly. "Once a model has done the hard work of learning video understanding, it will learn the causal relationship between video and audio to predict speech synchronized to lip movement and audio effects synchronized to the physical events causing them" (BFL x mimic). In plain terms, if the model already knows a hammer is about to strike, generating the thud is cheap.

Consider what that means for a buyer like Devon, a motion designer at a 40-person agency who budgets roughly 3 hours of manual sound work onto every 15-second clip his team ships, syncing footsteps and door slams by hand in a separate editor. The promise here is not just video with sound. It is video where the audio was predicted from the same physics that drove the picture, so the slam lands on the exact frame the door closes. Whether FLUX 3 delivers that consistently is unknown outside early access, but the architecture is built for it, and for Devon the payoff would show up in those 3 hours per clip, not in novelty.

What this resolves: FLUX 3's audio is a byproduct of learning video, by design, not a separate feature.

What it leaves unsolved: "Almost free" in compute does not mean "always correct" in output. Lip-sync and effect timing at 20 seconds are exactly the kind of thing that breaks under real prompts.


The Four Product Lines and When You Can Actually Use Them

FLUX 3 is not one product you can sign up for. It is four lines at four different stages of availability, and confusing them is the fastest way to be disappointed. The table below reflects BFL's own rollout statements.

Product line What it does Availability today
FLUX 3 Video Text/image-to-video, up to 20s, native audio Early access now, approval-gated, no public API
FLUX 3 Image The unified image generation and editing tier Early access "in the following weeks"
FLUX 3 Action / FLUX-mimic Video-action model for robotics Selected research and commercial partners only
FLUX 3 Dev Open-weight backbone for local deployment Later in 2026, no firm date

Sources: Black Forest Labs for the four lines and staging, and Decrypt for the open-weight Dev tier being "the only tier BFL plans to release for local use," due "later in 2026."

If your first instinct after reading a launch article was to open FLUX 3 and make something, this table is why you could not. Only Video is live, and even that is a waitlist, not a signup. We walk through the application process, the timeline, and the alternatives in the companion piece on how to access FLUX 3.

What this resolves: "Can I use FLUX 3" has four different answers depending on which line you mean.

What it leaves unsolved: Timelines beyond "weeks" and "later in 2026" are not committed. Treat any specific date you see elsewhere as a guess.

Waiting on FLUX 3 access? Ship content now.
You cannot generate FLUX 3 video yet, but you can point an agent at a trend and get video concepts, LinkedIn and X posts, headlines, and taglines in one pass. MoClaw researches live sources, then produces the content package for you.
Turn a trend into a full content package…Try MoClaw →

Correcting the Record: What the Pre-Launch Leaks Got Wrong

If you researched FLUX 3 before July 23, you probably read something false, and it is worth saying why.

A placeholder page appeared on BFL's own domain a couple of days before launch, and several aggregator sites raced to publish explainers based on it. Those pieces asserted two things that the actual launch contradicted: that FLUX 3 was unreleased with no date and no benchmarks, and that it was purely an image model competing in a different category than video generators. Both were wrong the moment the model shipped as a multimodal, video-first system with published preference numbers.

Think of Priya, a content lead who bookmarked one of those early "What is FLUX 3" pages on July 21, read that the model was unreleased and image-only, and crossed it off her Q3 planning doc. Two days later it shipped as a video-first model, and she had already written it off. She lost nothing but a place near the front of the early-access queue, which is exactly the cost of trusting a leaked page. The correction is simple: FLUX 3 released on July 23, 2026, it leads with video and audio, and it ships with benchmark claims. Any article telling you otherwise is describing the leaked placeholder, not the product.

The safe move is to date-check anything you read about FLUX 3 against the July 23 announcement, because a model that changed category between its leak and its release breaks every explainer written from the leak.

What this resolves: "Unreleased" and "image-only" are leak-era artifacts. Discard them.

What it leaves unsolved: Plenty of stale pages are still ranking. Expect to encounter the wrong version of this story for a while.


FLUX 3 Benchmarks Come With a Big Caveat

BFL published head-to-head preference numbers, and they deserve to be read carefully rather than repeated.

On BFL's own evaluation, FLUX 3 was preferred over Runway Gen-4.5 77% of the time and over Luma Ray 3.2 93% of the time. Against the models that matter most, it was preferred 52% of the time over both Seedance 2.0 and Gemini Omni Flash (Black Forest Labs). A 52% preference is a coin flip, not a win, and it is telling that the lopsided victories are against the weaker field while the toughest comparisons land on a tie.

FLUX 3 preference win rates from Black Forest Labs' own evaluation, which BFL labels a preliminary test of an early FLUX 3 candidate. Source: Black Forest Labs.
FLUX 3 preference win rates from Black Forest Labs' own evaluation, which BFL labels a preliminary test of an early FLUX 3 candidate. Source: Black Forest Labs.

Two caveats sit on every one of these numbers. First, they are BFL's own preference tests, not independent evaluation. Second, per the launch reporting, the chart measured a pre-release FLUX 3 candidate rather than the shipping model, so the numbers describe a checkpoint, not necessarily what you would get. Stack those together and the honest reading is: FLUX 3 is competitive at the top of the field, decisively ahead of the middle, and unproven by anyone but its maker.

We dig into what the Seedance comparison actually means for a buyer, including a wrinkle that makes the whole matchup stranger than the percentages suggest, in FLUX 3 vs Seedance 2.0.

What this resolves: The benchmark story is "competitive, self-reported, pre-release," not "new state of the art."

What it leaves unsolved: Independent, apples-to-apples testing does not exist yet, because most evaluators cannot get access.


FLUX 3 and the AI Agent Stack

Here is the part of the FLUX 3 story that is easy to miss under the video demos. The FLUX family is already wired into agent products, not just creative apps. BFL lists its models as "powering generative features inside leading creative, developer, and consumer platforms like Adobe Photoshop, Picsart, Nous Research's Hermes Agent, and more" (Yahoo Finance / GlobeNewswire).

That Hermes Agent detail matters, and so does BFL's own phrase for stitching clips: "agent-driven links." The direction is clear. Visual models like FLUX 3 are becoming tools that agents call, plan around, and orchestrate, rather than apps a human clicks through one render at a time. The interesting layer is not the model. It is the agent deciding what to generate, in what order, and what to do with the result.

That is the layer MoClaw works at. MoClaw is a managed AI agent platform, and while it does not run FLUX 3, the pattern FLUX's own customer list points to, agents orchestrating visual and content generation, is exactly what an agent is good for. If your real goal behind watching this launch is turning an idea into finished content on a schedule, the agent is the part you can use today.

What this resolves: FLUX 3 is already an agent-called model, and the orchestration layer is usable even when the model is not.

What it leaves unsolved: No public FLUX 3 API means no third-party platform, MoClaw included, can integrate it yet. That waits on BFL.


FAQ

Is FLUX 3 released?

Yes. FLUX 3 was announced on July 23, 2026, and FLUX 3 Video is in early access now. Any source calling it unreleased is describing a placeholder page that leaked before launch. Access is approval-gated, so "released" does not mean "open to everyone."

Is FLUX 3 open source?

Not yet. BFL plans an open-weight tier called FLUX 3 Dev for local deployment, but it is due later in 2026 with no firm date. The Video tier in early access today is not open weight.

Is FLUX 3 free?

BFL has not published any pricing for FLUX 3. Early access is an application, and the company has said nothing about cost tiers, so anyone quoting a FLUX 3 price is guessing.

What is the FLUX 3 release date?

July 23, 2026, for the announcement and FLUX 3 Video early access. FLUX 3 Image is slated for early access within weeks of launch, and the open-weight Dev tier for later in 2026.

FLUX 3 vs FLUX.2, what changed?

FLUX.2 (November 2025) was an image model with a lukewarm reception, and Alibaba's Z-Image Turbo took the open-source crown in late 2025 (Decrypt). FLUX 3 is the category jump: a unified multimodal model that leads with video and native audio rather than still images.


What FLUX 3 Means Before You Can Touch It

The honest one-line summary is that FLUX 3 is an ambitious multimodal bet you mostly cannot run yet. The architecture is genuinely new, the native-audio-from-video idea is elegant, and the numbers are competitive at the top of the field. All of that is real. It is also true that only one of four product lines is live, that line is gated, there is no public API, and the benchmarks are self-reported on a pre-release checkpoint.

So treat this launch the way you would treat any pre-access model: understand it, date-check what you read, and get on the waitlist if the video capability fits your work. Then keep shipping with what you can actually use. If your aim is finished content rather than a specific model, an agent that researches a trend and produces the whole package is available now, and it will be ready to reach for FLUX 3 the day a public API exists. For the how-and-when of getting in, start with how to access FLUX 3.

Continue Reading

M
MoClaw Editorial MoClaw editorial team

The MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.

Ready to put this into practice?

MoClaw runs browser tasks, research, and schedules automatically. Try it free.

flux 3 flux 3 multimodal flux 3 explained flux 3 release date flux 3 vs flux.2 flux 3 video

References: FLUX 3: Towards Multimodal Flow Models as the Backbone of Visual Intelligence (Black Forest Labs) · FLUX 3 x mimic: video-action models and the compute behind FLUX 3 (Black Forest Labs) · FLUX 3 generates videos with native audio up to 20 seconds long (The Decoder) · Black Forest Labs' FLUX 3 does image, video, and robot hands (Decrypt) · Black Forest Labs unveils FLUX 3 (Yahoo Finance / GlobeNewswire)