Black Forest Labs has introduced FLUX 3, a new multimodal foundation model that jointly learns from images, video, and audio within a single unified architecture. The company announced the model on July 23, 2026, making it available immediately through an early access program for developers and researchers. FLUX 3 represents a departure from modality-specific approaches by treating images, video, and audio as different projections of the same underlying reality, each captured by distinct sensors that lose different information in the process.
The model builds on the company's Self-Flow approach, which aligns multimodal generation and understanding within the same underlying architecture. By significantly scaling compute and data resources, the company trained FLUX 3 across all three modalities simultaneously. Early evaluations indicate strong performance across video generation with native audio, image synthesis and editing, and action prediction for robotics applications. The company also revealed a partnership with mimic robotics to develop FLUX-mimic, a specialized video-action model for dexterous manipulation tasks.
What's New / Specs
FLUX 3 introduces several capabilities that distinguish it from previous generation models. The video generation component can create clips up to 20 seconds in length with synchronized audio from text prompts, image references, or existing video clips. The system supports multiple generation modes including text-to-video, image-to-video animation, video-to-video transformation, generative video-audio continuation, and keyframe-to-video transitions. Multilingual dialogue generation and agentic chaining of clips into multi-shot sequences extend the creative possibilities.
- Video generation up to 20 seconds with native audio synthesis
- Text-to-video, image-to-video, video-to-video, and keyframe-to-video modes
- Multilingual dialogue and typography generation
- Agentic chaining for multi-shot sequences lasting several minutes
- Image synthesis and editing across styles, aspect ratios, and resolutions
- High-accuracy multilingual text rendering in images
- Action prediction via integrated approach and FLUX-mimic robotics partnership
- Open-weight Dev version planned for content creation and action prediction
Early benchmark comparisons published by the company show FLUX 3 Video preferred over several competing systems in human evaluation. The model was favored over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. More decisively, FLUX 3 led Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%. The company emphasizes these are preliminary results from 10-second 720p text-to-video clips with audio, and further improvements are expected during the early access phase.
The image generation component demonstrates significant improvement over earlier FLUX versions in handling complex prompts and text generation. The company reports the model produces a wide range of output styles and renders high-accuracy text in multiple languages. Early access for FLUX 3 Image will open in the following weeks. For action prediction, the company has pursued two routes: integrating native action prediction directly into FLUX 3, and using the pretrained video backbone as a dynamics-aware foundation for specialized action models fine-tuned with limited task-specific data.
The Self-Flow methodology underpinning FLUX 3 is presented as an evolution of flow matching. According to the company's own analysis, Self-Flow achieves lower generation error measured by Fréchet distance across modalities and higher success rates on manipulation tasks after fine-tuning compared with standard Flow Matching. The training run for FLUX 3 scaled this approach with substantially increased compute and data, enabling joint learning across video, images, and audio rather than sequential or separate training pipelines.
Why It Matters
The unified multimodal approach addresses a fundamental limitation of modality-specific models. Images capture spatial structures at a single moment, video restores temporal dynamics and physical laws, and audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. By training on all modalities simultaneously, FLUX 3 learns mutual constraints: the sound must match the impact, the motion must obey mass, and the future must follow from the past. This creates a more coherent world model suitable for both content creation and physical AI applications.
The partnership with mimic robotics illustrates the practical implications for physical AI. FLUX-mimic combines the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation and production deployment. The collaboration is already being tested on real production tasks at Audi, demonstrating how the same foundation model can serve creative content generation and industrial robotics. This dual-use capability suggests a convergence path where perceptual understanding, action prediction, and language processing share a common representation.
The staged launch plan reflects a deliberate approach to safety and reliability. The company will roll out video and audio generation through APIs and private weight access first, followed by action prediction through selected research and commercial partners, then image synthesis and editing, and finally an open-weight Dev version. Each capability undergoes an early access phase for feedback collection and rigorous safety testing. The company also plans to release more technical details on the underlying Self-Flow approach, which could inform broader research on multimodal flow matching architectures.
Beyond the immediate product lineup, the announcement signals a strategic bet that a single multimodal backbone can serve both generative media and embodied intelligence. The ability to chain clips agentically into multi-minute sequences while preserving character consistency across scenes hints at future workflows for long-form video production. Similarly, the claim that the video backbone can be fine-tuned for manipulation with limited data suggests a path to reduce the data hunger that has historically constrained robot learning.
Our Take
FLUX 3 represents a meaningful step toward unified multimodal intelligence, but several caveats deserve attention. The early access designation and preliminary evaluation metrics indicate the model is still maturing. Human preference benchmarks, while encouraging, cover specific clip lengths and resolutions that may not generalize to all use cases. The 93% preference over Luma Ray 3.2 and 77% over Runway Gen-4.5 are notable, but the comparison set excludes some frontier models like Sora or Veo, making the competitive positioning incomplete.
The robotics integration via FLUX-mimic is perhaps the most consequential development. Using a video generation backbone as a dynamics-aware foundation for action models is an elegant approach that could reduce the data requirements for robot learning. However, the gap between simulation-aware video prediction and real-world closed-loop control remains substantial. The Audi production testing will provide valuable signal on whether this approach transfers to industrial deployment. The open-weight Dev version, when released, will be a critical test of whether the community can extend and audit the model's capabilities independently.
The company's decision to keep the initial rollout behind APIs and private weight access rather than immediate open release reflects a cautious stance on misuse potential, especially given the model's ability to generate synchronized audio-visual content. The promise of later open-weight access for a Dev variant suggests a tiered strategy similar to other frontier labs, but the timeline remains unspecified. Until the safety testing methodology and red-teaming results are published, external assessment of risk mitigation will be limited.
FAQ
What modalities does FLUX 3 support and what are the output limits?
FLUX 3 jointly processes images, video, and audio within a single architecture. Video generation supports clips up to 20 seconds with native audio synthesis. Image synthesis handles multiple styles, aspect ratios, and resolutions with high-accuracy multilingual text rendering. Action prediction is available through integrated capabilities and the FLUX-mimic specialized model.
How does FLUX 3 compare to other video generation models in early testing?
In preliminary human evaluations on 10-second 720p text-to-video clips with audio, FLUX 3 was preferred over Grok Imagine Video (69%), Kling v3 Pro (60%), Happy Horse v1 (59%), Happy Horse 1.1 (57%), Seedance 2.0 and Gemini Omni Flash (52%), Runway Gen-4.5 (77%), and Luma Ray 3.2 (93%). The company emphasizes these are early results and further improvements are expected.
What is the Self-Flow approach and how does it differ from standard flow matching?
Self-Flow is the company's method for efficiently aligning multimodal generation and understanding within the same underlying architecture. The company reports it achieves lower generation error (measured by Fréchet distance) across modalities and higher success rates on manipulation tasks through fine-tuning compared to standard Flow Matching. FLUX 3 scales this approach with significantly increased compute and data.
When will the different FLUX 3 capabilities become available?
FLUX 3 Video is available in early access immediately. FLUX 3 Image early access opens in the following weeks. Action prediction through FLUX-mimic and FLUX 3 Action is available to selected research and commercial partners beginning with mimic robotics. The open-weight FLUX 3 Dev version will be released after early access phases for each capability, with no specific date announced.
What safety and testing measures are in place for the rollout?
Each capability undergoes an early access phase for smooth rollout, feedback collection, and rigorous safety testing before broader availability. The staged release sequence — video first, then action prediction, then image, then open-weight Dev — allows iterative validation. The company has not published specific safety benchmarks or red-teaming details in the announcement.