What Is FLUX 3? Capabilities, Early Access, and Release Status
A fact-checked guide to FLUX 3, Black Forest Labs’ multimodal model for image, video, native audio, and action, including confirmed features and what remains unknown.
FLUX 3 is Black Forest Labs' new multimodal model family for generating and understanding images, video, audio, and action within one shared system. It is broader than an image-model update: BFL describes a foundation trained across visual and audio data, with video generation, image synthesis and editing, and action prediction released through separate stages.
The important availability detail is easy to miss. FLUX 3 is not generally available, but it is no longer only a rumor or placeholder. Black Forest Labs announced FLUX 3 officially on July 23, 2026 and opened limited Early Access for FLUX 3 Video. Image Early Access is planned for the following weeks. General release dates, public API prices, and complete technical specifications remain unpublished.
MuseFable does not currently offer FLUX 3 generation. You can review the live status and join the FLUX 3 waitlist with your account, then receive one email when the model becomes available here.
What makes FLUX 3 different?
Earlier FLUX releases were best known for image generation and editing. FLUX 3 expands that scope by learning from images, video, and audio together. BFL says this shared training lets the model mix input and output modalities instead of routing every task through an isolated model.
That matters for workflows where one medium depends on another. A reference image can guide a video. A source video can carry a character into a new scene. Sound can be generated with the visual sequence, rather than added later without knowledge of what happens on screen.
BFL connects the model to its Self-Flow research, an approach intended to align multimodal generation and understanding within one architecture. The company has not yet published the full FLUX 3 model card, training details, parameter count, or inference requirements, so claims beyond the announcement should be treated as provisional.
Confirmed FLUX 3 Video capabilities
FLUX 3 Video is the first branch in Early Access. According to BFL, it supports:
- Text-to-video generation with native audio.
- Image-to-video from a starting frame or visual reference.
- Video-to-video using a source clip as a reference.
- Video and audio continuation from an existing sequence.
- Keyframe-to-video transitions between defined moments.
- Multilingual dialogue and a wide range of visual styles.
- Clips up to 20 seconds in one generation.
- Chaining individual clips into longer, multi-shot sequences.
BFL's preliminary evaluations used 10-second, 720p clips with audio. The company reports favorable preference results against several current video models, but also says both the model and its evaluation harness are still being improved. These are vendor-reported early results, not independent benchmarks.
Image, Action, and Dev releases
FLUX 3 Image is intended for image synthesis and editing across styles, aspect ratios, and resolutions. BFL highlights improved handling of complex prompts and multilingual typography compared with earlier FLUX versions. Image Early Access is planned after the initial Video rollout.
FLUX 3 Action applies the same world-model backbone to action prediction. BFL is developing both native action prediction and specialized models derived from the video backbone. Its first named partner is mimic robotics. This branch is aimed at research and selected commercial partners, not general creative-tool access today.
FLUX 3 Dev is the planned open-weight multimodal backbone. BFL has named it in the launch plan, but has not published the license, weights, hardware requirements, or release date. Do not assume that licensing will match an earlier FLUX Dev release until BFL publishes the terms.
Is FLUX 3 released?
The accurate answer is: partially, through limited Early Access.
- FLUX 3 Video: available through BFL Early Access for selected users.
- FLUX 3 Image: Early Access planned for the following weeks.
- FLUX 3 Action: planned for selected research and commercial partners.
- FLUX 3 Dev: announced as a future open-weight multimodal backbone.
- Public pricing and general API access: not announced.
An earlier Kie.ai overview captured the model's initial placeholder-page period and the open questions surrounding it. It was published before BFL's full announcement, so its statements that no official release or waitlist existed were superseded later on July 23. It remains useful as a record of the pre-announcement signals, not as the current availability source.
How does FLUX 3 compare with FLUX.2?
FLUX.2 focuses on image generation and editing. FLUX 3 keeps image work but adds video, native audio, and action prediction to a shared multimodal foundation. The larger change is architectural scope: FLUX.2 is an image model family, while FLUX 3 is positioned as a model of visual and physical sequences.
That does not automatically make FLUX 3 the right production choice today. FLUX.2 and other released models have known endpoints, prices, and operating constraints. FLUX 3 still has gated access and incomplete public specifications.
What can you use while waiting?
If your goal is to produce images or video now, use a released model and keep the workflow portable:
- Create images with GPT Image 2 for prompt-based image production and editing.
- Create video with Seedance 2.0 Fast for faster video drafts.
- Build a Product Content Pack that combines image, video, voice, and copy.
- Turn a script into an AI explainer video.
These tools let you validate prompts, references, brand rules, and review steps before FLUX 3 becomes broadly available. When FLUX 3 is ready for reliable use here, the model can enter an existing workflow instead of forcing you to invent one around an unreleased specification.
Frequently asked questions
Can I use FLUX 3 here now?
Not yet. MuseFable will mark the model available only after reliable access, pricing, and operating constraints are known. You can join the FLUX 3 waitlist now.
Does FLUX 3 generate audio?
Yes. BFL says FLUX 3 Video generates native audio with video, including multilingual dialogue and sounds associated with visual events.
How long are FLUX 3 videos?
BFL says the model can generate video with audio up to 20 seconds in a single generation. Its published preliminary evaluation used 10-second, 720p outputs.
Is FLUX 3 open source?
BFL has announced a future FLUX 3 Dev open-weight backbone. The weights, license, release date, and hardware requirements have not been published.
How much will FLUX 3 cost?
Public API pricing has not been announced. Any price presented before BFL or a supported provider publishes an official rate is speculation.