What Is FLUX 3? BFL's Multimodal Omni Model

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: July 23, 2026
Placeholder page for FLUX 3 on the Black Forest Labs site describing a multimodal image, video, audio, and action model

TLDRFLUX 3 is Black Forest Labs' unreleased multimodal model generating image, video, audio, and action from one network. Confirmed facts, open questions, and access.

FLUX 3 Is Black Forest Labs' Omni Model — One Network for Image, Video, Audio, and Action

FLUX 3 is an unreleased multimodal generative model from Black Forest Labs (BFL) built to produce image, video, audio, and action outputs from a single network. It first surfaced on July 21, 2026 via a placeholder page at bfl.ai/models/flux-3 that was live briefly and then pulled down. The page carried one tagline: "A breakthrough in control, realism, and world understanding — one multimodal model generating image, video, audio and action." No release date, pricing, parameter count, or benchmark has been published by BFL as of July 23, 2026.

Key Takeaways

  • FLUX 3 is Black Forest Labs' next-generation "omni" model, positioned as a single network for image, video, audio, and action generation.
  • The model was first spotted on July 21, 2026 via a placeholder page at bfl.ai/models/flux-3 that was quickly removed; an archived copy exists on the Wayback Machine.
  • BFL CEO Robin Rombach (@robrombach) posted an uncaptioned ~20-second video clip on July 21, 2026, widely read as a FLUX 3 teaser; the official @bfl_ai account replied with a single "👀" emoji.
  • Select users appear to have early access and are posting generations without naming the model.
  • Open weights status, pricing, resolution, context length, and release date are all unconfirmed.
  • FLUX 3 succeeds the released FLUX.2 family (a text-to-image and image-to-image model) with a much broader modality footprint.

What Is FLUX 3?

FLUX 3 is Black Forest Labs' upcoming multimodal foundation model. Where earlier FLUX releases focused on image generation and editing, FLUX 3 is described on BFL's own (now removed) placeholder page as a single network that generates image, video, audio, and action. That framing places it in the "omni model" category alongside efforts from OpenAI, Google, and ByteDance, rather than in the narrower text-to-image lane.

The model has not been officially released. What exists publicly is a brief placeholder page, a CEO teaser clip, and a small stream of early-access generations circulating without attribution. Andrew Curran, who tracks BFL closely, summarized the state on July 23, 2026: "Black Forest Labs video model, which I posted about when it was first teased on their site back in August 2024, is about to arrive in the form of a multimodal Flux 3." source

Black Forest Labs video model, which I posted about when it was first teased on their site back in A

Source: @AndrewCurran_

Black Forest Labs is the Berlin-based lab founded by former Stability AI researchers, including CEO Robin Rombach. The FLUX family shipped in 2024 with FLUX.1 [pro], [dev], and [schnell], expanded in 2025 with the FLUX.1 Kontext image editing line, and reached FLUX.2 in 2026. FLUX 3 is the first BFL model to publicly claim video and audio generation.

FLUX 3 at a Glance

AttributeValue
DeveloperBlack Forest Labs (BFL)
TypeMultimodal generative model ("omni")
ModalityImage, video, audio, action (per placeholder page)
Context / inputReported to accept image and video inputs; not confirmed
Video lengthDemo clip ~20 seconds; official specs not yet confirmed
AudioReportedly native audio (dialogue, SFX, music); unconfirmed
ResolutionNot yet confirmed
PricingNot yet confirmed
AvailabilityLimited early access; no public release
License / open weightsNot yet confirmed
Official pagebfl.ai/models/flux-3 (currently 404)

How FLUX 3 Works and What Makes It Different

The one confirmed capability line is BFL's own tagline: "A breakthrough in control, realism, and world understanding — one multimodal model generating image, video, audio and action." That single sentence carries the entire product thesis, so it is worth reading closely.

Three coined anchors follow from it. World Understanding signals that FLUX 3 is being pitched as more than a pixel generator — the phrase is closer to how DeepMind and NVIDIA describe world models than how BFL described earlier FLUX releases. Multimodal Omni Output means one shared network produces four output types instead of a specialist per modality. Action Generation is the most unusual claim; it hints at motion planning or control signals rather than only rendered frames, though BFL has not clarified whether "action" means robotics-grade control, in-scene character motion, or something else.

Community observers have reported additional specifics based on unofficial channels — up to 4MP images, 10+ reference images, 1080p 24fps video, 48kHz stereo audio, and "real-time world-state prediction" — but these numbers appear in a community-collected list, not a BFL document, and should be treated as unconfirmed. Per an r/StableDiffusion thread on the placeholder page, even the source of that list is uncertain.

The direction is coherent with BFL's public research bets since 2024, when the lab first teased a video model. FLUX 3 appears to be that video effort landing inside a broader multimodal wrapper rather than as a standalone product.

What You Can Do With FLUX 3

Confirmed use cases are limited to what the tagline implies: single-prompt generation of any of image, video, audio, or action, presumably with cross-modal conditioning (e.g., image plus text into video with audio). The BFL CEO's ~20-second teaser clip is the only public sample tied directly to the model. Community discussion of unlabeled early-access outputs is circulating but cannot be treated as a spec sheet.

For teams already building on FLUX for stills, a natural adjacent workflow today is generating base images with a current-generation model and editing them iteratively. On kie.ai you can try FLUX.2 for that image-side work while FLUX 3's video and audio capabilities remain gated.

How FLUX 3 Compares

FLUX 3 sits in a crowded lane. The direct comparison set for a multimodal image-plus-video model includes ByteDance's Seedance line, Kuaishou's Kling, Google's Veo, and OpenAI's Sora — with the added twist that FLUX 3 also claims audio and action.

ModelVendorModalitiesOpen WeightsStatus
FLUX 3Black Forest LabsImage, video, audio, actionNot yet confirmedUnreleased
FLUX.2Black Forest LabsImage (T2I, I2I)Partial (Dev tier)Released
Seedance 2.5ByteDanceVideo (T2V, I2V, V2V)NoReleased
Veo 3.1GoogleVideo with audioNoReleased

The wild card is open weights. BFL has historically shipped a proprietary Pro tier alongside an openly licensed Dev tier. Whether FLUX 3 follows that pattern is the single most-asked question in community threads, and BFL has not answered it.

Availability: How to Access FLUX 3

FLUX 3 is not generally available. As of July 23, 2026:

  • The official model page at bfl.ai/models/flux-3 returns a 404.
  • An archived snapshot of the placeholder page is preserved at web.archive.org for the July 21, 2026 capture.
  • Select users appear to have received early access and are posting generations, per Andrew Curran's timeline observation, but participants have reportedly been told not to identify the model.
  • No API endpoint, pricing tier, waitlist form, or Hugging Face repository has been officially announced.

The authoritative channels to watch are bfl.ai, the @bfl_ai account, and CEO Robin Rombach's @robrombach account. Any URL containing "flux-3.com" or similarly branded consumer sites is not affiliated with Black Forest Labs.

What We Don't Know Yet

For a reference page, the honest gaps matter as much as the confirmed facts. Open questions as of July 23, 2026:

  • Release date. No public window.
  • Pricing. No API tier, subscription, or per-generation cost announced.
  • Open weights. Whether BFL will publish a Dev or Klein variant, and under what license.
  • Parameter count and hardware requirements. Nothing disclosed.
  • Video specs. Native resolution, maximum duration, frame rate, and audio sync quality are unconfirmed.
  • What "action" means. Robotics control, in-scene motion planning, or something else entirely is unclear.
  • Benchmarks. No official evaluations against Seedance, Veo, Kling, or Sora.
  • Early-access terms. Who has it, and under what NDA, is not public.

Frequently Asked Questions

What is FLUX 3?

FLUX 3 is an unreleased multimodal generative model from Black Forest Labs (BFL) designed to produce image, video, audio, and action outputs from a single network. It was first surfaced on July 21, 2026 via a placeholder page at bfl.ai/models/flux-3 that was quickly removed.

Is FLUX 3 released yet?

FLUX 3 has not been officially released as of July 23, 2026. Only a placeholder page, a CEO teaser video, and unlabeled generations from early-access users have surfaced. BFL has published no release date, specs, or pricing.

Is FLUX 3 open source?

Open weights status for FLUX 3 has not been confirmed. Black Forest Labs has historically shipped both proprietary Pro tiers and open-weight Dev variants for the FLUX.1 family, but there is no official statement yet on whether FLUX 3 will follow that pattern.

What can FLUX 3 do?

According to the removed placeholder page, FLUX 3 is one multimodal model generating image, video, audio, and action. BFL positions it as a breakthrough in control, realism, and world understanding. Community reports mention native audio in video and image plus video inputs, but these details are unconfirmed.

How much does FLUX 3 cost?

FLUX 3 pricing has not been published. No API tier, subscription plan, or per-generation cost has been announced by Black Forest Labs.

How is FLUX 3 different from FLUX.2?

FLUX.2 is a released text-to-image and image-to-image model from Black Forest Labs. FLUX 3 is positioned as a fully multimodal successor that adds video, audio, and action generation within a single network, rather than a specialized image model.

When will FLUX 3 be available?

No release date has been announced. The placeholder page was removed shortly after appearing on July 21, 2026, and only select users appear to have early access. Watch bfl.ai and the @bfl_ai account for the official launch.

What to Watch Next

Three signals will move this page from placeholder-analysis to full spec sheet. First, whether BFL re-publishes the bfl.ai/models/flux-3 URL with an official model card and pricing. Second, whether a Dev or Klein open-weights tier appears on Hugging Face, which would determine whether local inference is on the table. Third, whether early-access users are cleared to post benchmarks against Seedance, Veo, and Kling — those numbers, not the tagline, will settle where FLUX 3 actually lands.

Building similar multimodal image, video, and editing workflows today? On kie.ai you can try FLUX.2, Seedance 2.5, and Veo 3.1.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair