Black Forest Labs Releases 'FLUX 3': The 'One Multimodal Model' for Images, Video, Audio, and Robot Actions

In the span of a few years, generative models have moved from producing still images based on text descriptions to systems that attempt to capture motion, sound, and even the physical consequences of events. 

Most of the large technology companies have approached this by developing specialized systems for each modality and then connecting them through interfaces or post-processing steps. The result is often a collection of models that work together rather than a single representation of the world.

Black Forest Labs sits somewhat apart from that pattern. 

The company was founded in 2024 by a group of researchers who had previously worked on latent diffusion and the Stable Diffusion family of models at Stability AI.

Their earlier FLUX models focused on high-quality image generation and editing, released in both open-weight and commercial forms, and quickly became widely used for their prompt adherence and visual detail. 

Now, the company introduced 'FLUX 3.'

In a blog post, Black Forest Labs described FLUX 3 as a single multimodal model trained jointly on images, video, and audio. 

The architecture, which they call Self-Flow, is designed so that the constraints of one modality help shape the others. 

Spatial structure from still frames, temporal dynamics from video, and causal cues from sound are learned together rather than assembled afterward. Early access has been opened for video generation that includes native audio, with clips lasting up to twenty seconds. 

The same backbone is also being extended to action prediction. 

Image
FLUX 3
FLUX 3 expands on Self-Flow, which is a unified architecture for aligning multimodal generation and understanding. By scaling up both compute and training data, Black Forest Labs trained FLUX 3 across video, images, and audio simultaneously

What sets the approach apart from many larger efforts is the insistence on a shared foundation. 

Where some organizations maintain separate video models, audio models, and robot-control systems, FLUX 3 treats content creation and physical action as two applications of the same learned representation of how the world behaves. 

Video prediction forms the bulk of the training compute, giving the model an internal sense of contact, motion, and cause and effect. 

Action signals and audio are treated as additional low-dimensional views of that same reality. 

The company reports that adding the action pathway produces only a temporary dip in generation quality before recovery, suggesting the capacity cost is limited.

At the same time, the release arrives with clear constraints. 

Video generation is limited to roughly 20 seconds in a single pass, though clips can be chained. Image generation and editing remain unavailable to the public and are scheduled for a later early-access phase. 

Action prediction is restricted to selected research and commercial partners rather than open use. 

No public API, pricing, full technical report, or downloadable weights have been released yet, and the published preference evaluations are described as preliminary results on an early candidate rather than the final shipping model. 

Independent verification of longer-term consistency, higher resolutions, or edge-case reliability is therefore still limited.

The robotics work remains available through selected partners for now. 

Regardless, release places a clear bet on the idea that a model trained to understand the world as a coherent whole can serve both creative and physical applications without requiring separate foundations for each.

 

 

Published