TechBeetle | Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio - but in l...
Tech Beetle briefing US AI

Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio - but in limited release to start

Essential brief

Black Forest Labs (BFL) has introduced FLUX 3, a multimodal AI model capable of generating images, 20-second video clips with synchronized audio, and robotic action predictions from a single prompt

Key topics

black forest labs flux capable generating images 20-second video audio limited release start

Key facts

FLUX 3 generates images, 20-second videos with audio, and predicts robotic actions from one prompt.
The model uses a unified architecture jointly trained across video, audio, and images, unlike separate modality models.
Early access is gated; public API, pricing, and open-source weights will be available later in 2026.
FLUX 3 supports advanced video generation features including multilingual dialogue and multi-clip sequencing.

Highlights

FLUX 3 is Black Forest Labs' first public video generation model, producing clips up to 20 seconds with synchronized audio.
The model is available through early access for video and action products; image generation will follow soon.
Preliminary benchmarks show FLUX 3 preferred over several competitors in 10-second 720p video tests, but full evaluations are pending.
FLUX 3 Dev will provide open-weight access to a multimodal backbone for content creation and robotic action prediction later this year.
FLUX-mimic, developed with Mimic Robotics, uses FLUX 3 to improve robotic manipulation with significantly reduced training data requirements.

Why it matters

FLUX 3 represents a significant advancement in AI by integrating image, video, audio generation, and robotic action prediction within a single unified model. This multimodal approach could streamline creative workflows and enhance physical AI applications like robotics by reducing the need for separate models and extensive task-specific training. The model's development reflects broader industry trends toward unified architectures that better understand and interact with dynamic, multimodal environments.

Black Forest Labs (BFL), based in Freiburg, Germany, has expanded its FLUX AI model family with the launch of FLUX 3, a multimodal frontier model designed to generate images, video clips up to 20 seconds with synchronized audio, and predict robotic actions from a single prompt. Unlike assembling separate models for each modality, FLUX 3 is jointly trained across video, audio, and images within a unified architecture. This approach supports BFL's vision of "visual intelligence"—models that perceive, predict, and act across both physical and digital environments.

FLUX 3 is offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the upcoming open-source FLUX 3 Dev. Currently, FLUX 3 Video and FLUX 3 Action are available via a gated early access program requiring approval, with no public API or partner access yet. FLUX 3 Image is expected to roll out in the coming weeks, followed by general availability. Pricing, service-level agreements, and detailed benchmarks have not been disclosed.

Preliminary benchmark tests on 10-second, 720p text-to-video clips with audio show FLUX 3 outperforming competitors such as Luma Ray 3.2 and Runway Gen-4.5 in user preference tests. However, these results are based on early model versions and lack comprehensive evaluation methodologies. Notably, FLUX 3 does not launch with downloadable weights or an open-source license; these are planned for later releases, including FLUX 3 Dev, which will provide open-weight access to a multimodal backbone for content creation and action prediction.

The video generation capabilities include text-to-video, image-to-video, video-to-video, generative continuation, keyframe-to-video, multilingual dialogue, and animated design. FLUX 3 supports sequences lasting several minutes through agentic chaining of clips, addressing continuity challenges in generative video. The model is already being tested by companies like Canva, Burda, Magnific, Krea, and Picsart.

In robotics, FLUX 3 underpins FLUX-mimic, developed with Swiss firm Mimic Robotics, enabling robots to understand scenes, predict action outcomes, and adapt to new tasks with significantly less data than prior methods. This unified architecture approach contrasts with models trained solely on images, emphasizing the importance of multimodal training for physical and digital intelligence.

BFL, founded by key contributors to Stable Diffusion and other AI models, has built a reputation for open-source image generation models widely adopted in the industry. FLUX 3 Dev represents a significant expansion into multimodal and physical AI applications, with open weights expected to support secure, low-latency local deployment for enterprises. The company is valued at $3.25 billion and has raised over $450 million from investors including a16z, Salesforce Ventures, Nvidia, Adobe Ventures, and Deutsche Telekom's T.Capital.

Key topics in this update include black forest labs, flux, and capable.