What Is FLUX 3? BFL's Multimodal Omni Model

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: July 23, 2026
Placeholder page for FLUX 3 on the Black Forest Labs site describing a multimodal image, video, audio, and action model

TLDRFLUX 3 is Black Forest Labs' multimodal model generating image, video, audio, and action-prediction from one network. Launched July 23, 2026 with video early access open.

FLUX 3 Is Black Forest Labs' Omni Model — One Network for Image, Video, Audio, and Action

FLUX 3 is a multimodal generative model from Black Forest Labs (BFL) that produces image, video, audio, and action-prediction outputs from a single network. Black Forest Labs officially launched it on July 23, 2026, two days after a brief placeholder page at bfl.ai/models/flux-3 surfaced on July 21. The launch tagline: "A breakthrough in control, realism, and world understanding — one multimodal model generating image, video, audio and action." FLUX 3 Video is open now via gated early access; image generation follows in the coming weeks, and an open-weight FLUX 3 Dev backbone is planned for later.

Key Takeaways

  • FLUX 3 is Black Forest Labs' "omni" model, launched July 23, 2026 as a single network for image, video, audio, and action-prediction generation.
  • Video is the flagship capability at launch: up to 20 seconds per generation with native audio (dialogue, SFX, music), plus text-to-video, image-to-video, video-to-video, keyframe control, and multi-shot chaining.
  • FLUX 3 Video early access is open now via bfl.ai/models/flux-3 (application required); image early access is coming in the following weeks.
  • An open-weight FLUX 3 Dev multimodal backbone has been confirmed for later release; pricing and public API details are not yet published.
  • Action-prediction ships through partners first — starting with mimic robotics, which is using FLUX 3 as the backbone for a video-action model tested at Audi.
  • FLUX 3 succeeds the released FLUX.2 family (a text-to-image and image-to-image model) with a much broader modality footprint.

What Is FLUX 3?

FLUX 3 is Black Forest Labs' multimodal foundation model. Where earlier FLUX releases focused on image generation and editing, FLUX 3 is a single network that generates image, video, audio, and action-prediction. That framing places it in the "omni model" category alongside efforts from OpenAI, Google, and ByteDance, rather than in the narrower text-to-image lane.

The model is now live. On July 23, 2026, BFL published the official announcement, opened FLUX 3 Video early access, and unveiled a companion robotics collaboration. Andrew Curran, who tracks BFL closely, previewed the launch that morning: "Black Forest Labs video model, which I posted about when it was first teased on their site back in August 2024, is about to arrive in the form of a multimodal Flux 3." source

Black Forest Labs video model, which I posted about when it was first teased on their site back in A

Source: @AndrewCurran_

Black Forest Labs is the Berlin-based lab founded by former Stability AI researchers, including CEO Robin Rombach. The FLUX family shipped in 2024 with FLUX.1 [pro], [dev], and [schnell], expanded in 2025 with the FLUX.1 Kontext image editing line, and reached FLUX.2 in 2026. FLUX 3 is the first BFL model to ship with video and audio generation.

FLUX 3 at a Glance

AttributeValue
DeveloperBlack Forest Labs (BFL)
TypeMultimodal generative model ("omni")
ModalityImage, video, audio, action-prediction
Context / inputText, image, and video inputs (image refs, start frames, video refs)
Video lengthUp to 20 seconds per generation
AudioNative audio in video: dialogue, SFX, and music
ResolutionNot yet fully disclosed
PricingNot yet published
AvailabilityVideo early access open now; image early access in following weeks
License / open weightsOpen-weight FLUX 3 Dev backbone confirmed for later release
Official pagebfl.ai/models/flux-3

How FLUX 3 Works and What Makes It Different

The core product thesis comes straight from BFL's launch tagline: "A breakthrough in control, realism, and world understanding — one multimodal model generating image, video, audio and action." That one sentence carries the whole positioning.

Three coined anchors follow from it. World Understanding signals that FLUX 3 is being pitched as more than a pixel generator — the phrase is closer to how DeepMind and NVIDIA describe world models than how BFL described earlier FLUX releases. Multimodal Omni Output means one shared network produces four output types instead of a specialist per modality. Action-Prediction is the most unusual claim: BFL positions the FLUX 3 backbone as the foundation for a video-action model, and its launch partner mimic robotics is applying it to dexterous manipulation tasks — including tests at Audi — with reports of substantial gains in sample efficiency over standard vision-language-action models.

The video feature set confirmed at launch is broad: text-to-video and image-to-video (with images used as references or start frames), video-to-video from a reference clip, video and audio extension, keyframe-based generation, multilingual dialogue, a wide range of visual styles and aspect ratios, "smart editing" for chaining single clips into longer multi-shot sequences, and text-and-animation generation. It also outputs images, not just video. Per an r/StableDiffusion thread on the launch, early-access testers have been posting text-to-video and image-to-video generations with strong physics and cinematography.

The direction is coherent with BFL's public research bets since 2024, when the lab first teased a video model. FLUX 3 is that video effort landing inside a broader multimodal wrapper.

What You Can Do With FLUX 3

Confirmed use cases include single-prompt generation across image, video (up to 20 seconds with synced audio), and multi-shot sequences, with cross-modal conditioning — for example, image plus text into video with audio, or a reference clip into a new video with a preserved character. Early-access users are posting cinematic action sequences, character-driven scenes, and hyper-realistic tests: Mark Kretschmann shared a 20-second Godzilla clip generated with in-sync audio, and Umesh's continuous-shot rally-car sequence has been widely circulated.

For teams already building on FLUX for stills, a natural adjacent workflow today is generating base images with a current-generation model and editing them iteratively. On kie.ai you can try FLUX.2 for that image-side work while FLUX 3 video access is rolling out.

How FLUX 3 Compares

FLUX 3 sits in a crowded lane. The direct comparison set for a multimodal image-plus-video model includes ByteDance's Seedance line, Kuaishou's Kling, Google's Veo, and OpenAI's Sora — with the added twist that FLUX 3 also ships audio and action-prediction.

ModelVendorModalitiesOpen WeightsStatus
FLUX 3Black Forest LabsImage, video, audio, actionDev backbone plannedLaunched (early access)
FLUX.2Black Forest LabsImage (T2I, I2I)Partial (Dev tier)Released
Seedance 2.5ByteDanceVideo (T2V, I2V, V2V)NoReleased
Veo 3.1GoogleVideo with audioNoReleased

The open-weights question is now partially answered. BFL has confirmed an open-weight FLUX 3 Dev multimodal backbone, in line with how the lab has historically shipped a proprietary tier alongside an openly licensed Dev tier for FLUX.1 and FLUX.2. Timing and license terms for FLUX 3 Dev have not been announced.

Availability: How to Access FLUX 3

FLUX 3 launched on July 23, 2026 with a phased rollout. As of July 24, 2026:

  • The official model page is live at bfl.ai/models/flux-3, with a "Request early access" flow for FLUX 3 Video.
  • Video early access is open now to approved applicants; image generation early access is coming in the following weeks.
  • Action-prediction is arriving through partners first, starting with mimic robotics (which has already deployed a FLUX 3–based system in tests at Audi).
  • The open-weight FLUX 3 Dev multimodal backbone is planned for later release. No public API, pricing tier, or Hugging Face repository has been announced yet; BFL has indicated API access will open in the following weeks.

The authoritative channels to watch are bfl.ai, the @bfl_ai account, and CEO Robin Rombach's @robrombach account. Any URL containing "flux-3.com" or similarly branded consumer sites is not affiliated with Black Forest Labs.

What We Don't Know Yet

Even with the model launched, several details remain outstanding as of July 24, 2026:

  • Pricing. No API tier, subscription, or per-generation cost has been announced.
  • Public API timing. BFL has said API access is coming "in the following weeks" without a firm date.
  • Open-weight release. FLUX 3 Dev is confirmed but has no announced release date or license terms.
  • Parameter count and hardware requirements. Nothing officially disclosed.
  • Full video specs. Native resolution, maximum frame rate, and detailed audio sync specs beyond "native audio in video up to 20 seconds" are not yet published.
  • Formal benchmarks. No published head-to-head evaluations against Seedance, Veo, Kling, or Sora.

Frequently Asked Questions

What is FLUX 3?

FLUX 3 is a multimodal generative model from Black Forest Labs (BFL) that produces image, video, audio, and action-prediction outputs from a single network. It was officially launched on July 23, 2026, following a brief placeholder page at bfl.ai/models/flux-3 that appeared on July 21, 2026.

Is FLUX 3 released yet?

Yes. Black Forest Labs officially launched FLUX 3 on July 23, 2026. FLUX 3 Video is available now through a gated early-access program, image generation early access is coming in the following weeks, and an open-weight FLUX 3 Dev multimodal backbone is planned for release later.

Is FLUX 3 open source?

Black Forest Labs has confirmed that an open-weight version, FLUX 3 Dev, will be released later. As with the prior FLUX.1 and FLUX.2 Dev tiers, this is expected to use BFL's non-commercial open-weights license rather than a fully permissive open-source license. Exact timing has not been announced.

What can FLUX 3 do?

FLUX 3 generates images, video up to 20 seconds long with native audio (dialogue, SFX, music), and action-prediction from a single multimodal network. Video features include text-to-video, image-to-video, video-to-video, keyframe control, multi-shot chaining, multilingual dialogue, and a broad range of styles and aspect ratios.

How much does FLUX 3 cost?

FLUX 3 pricing has not been published. No API tier, subscription plan, or per-generation cost has been announced by Black Forest Labs at launch; the API is expected to open in the following weeks.

How is FLUX 3 different from FLUX.2?

FLUX.2 is a text-to-image and image-to-image model from Black Forest Labs. FLUX 3 is a fully multimodal successor that adds video generation (up to 20 seconds), native audio, and action-prediction within a single network, rather than a specialized image model.

When will FLUX 3 be available?

FLUX 3 launched on July 23, 2026. Video is available now via gated early access at bfl.ai/models/flux-3, image generation early access is rolling out in the following weeks, action-prediction is arriving through partners (starting with mimic robotics), and the open-weight FLUX 3 Dev backbone is planned for later in 2026.

What to Watch Next

Three signals will define how FLUX 3 lands. First, the public API and pricing — BFL has said access is coming in the following weeks, and those numbers will determine whether FLUX 3 is a serious challenger to Seedance and Veo for production video workloads. Second, the FLUX 3 Dev open-weight release: license terms, model size, and whether local inference of a full multimodal backbone is realistic will decide the on-prem and robotics story. Third, independent benchmarks against Seedance 2.5, Veo 3.1, Kling, and Sora — those, not the launch tagline, will settle where FLUX 3 actually sits.

Building similar multimodal image, video, and editing workflows today? On kie.ai you can try FLUX.2, Seedance 2.5, and Veo 3.1.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair