Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio - but in limited release to start
Essential brief
Black Forest Labs (BFL) has introduced FLUX 3, a multimodal AI model capable of generating images, 20-second video clips with synchronized audio, and robotic action predictions from a single prompt
Key topics
Key facts
Highlights
Why it matters
FLUX 3 represents a significant advancement in AI by integrating image, video, audio generation, and robotic action prediction within a single unified model. This multimodal approach could streamline creative workflows and enhance physical AI applications like robotics by reducing the need for separate models and extensive task-specific training. The model's development reflects broader industry trends toward unified architectures that better understand and interact with dynamic, multimodal environments.
Black Forest Labs (BFL), based in Freiburg, Germany, has expanded its FLUX AI model family with the launch of FLUX 3, a multimodal frontier model designed to generate images, video clips up to 20 seconds with synchronized audio, and predict robotic actions from a single prompt. Unlike assembling separate models for each modality, FLUX 3 is jointly trained across video, audio, and images within a unified architecture. This approach supports BFL's vision of "visual intelligence"—models that perceive, predict, and act across both physical and digital environments.
FLUX 3 is offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the upcoming open-source FLUX 3 Dev. Currently, FLUX 3 Video and FLUX 3 Action are available via a gated early access program requiring approval, with no public API or partner access yet. FLUX 3 Image is expected to roll out in the coming weeks, followed by general availability. Pricing, service-level agreements, and detailed benchmarks have not been disclosed.
Preliminary benchmark tests on 10-second, 720p text-to-video clips with audio show FLUX 3 outperforming competitors such as Luma Ray 3.2 and Runway Gen-4.5 in user preference tests. However, these results are based on early model versions and lack comprehensive evaluation methodologies. Notably, FLUX 3 does not launch with downloadable weights or an open-source license; these are planned for later releases, including FLUX 3 Dev, which will provide open-weight access to a multimodal backbone for content creation and action prediction.
The video generation capabilities include text-to-video, image-to-video, video-to-video, generative continuation, keyframe-to-video, multilingual dialogue, and animated design. FLUX 3 supports sequences lasting several minutes through agentic chaining of clips, addressing continuity challenges in generative video. The model is already being tested by companies like Canva, Burda, Magnific, Krea, and Picsart.
In robotics, FLUX 3 underpins FLUX-mimic, developed with Swiss firm Mimic Robotics, enabling robots to understand scenes, predict action outcomes, and adapt to new tasks with significantly less data than prior methods. This unified architecture approach contrasts with models trained solely on images, emphasizing the importance of multimodal training for physical and digital intelligence.
BFL, founded by key contributors to Stable Diffusion and other AI models, has built a reputation for open-source image generation models widely adopted in the industry. FLUX 3 Dev represents a significant expansion into multimodal and physical AI applications, with open weights expected to support secure, low-latency local deployment for enterprises. The company is valued at $3.25 billion and has raised over $450 million from investors including a16z, Salesforce Ventures, Nvidia, Adobe Ventures, and Deutsche Telekom's T.Capital.
Key topics in this update include black forest labs, flux, and capable.