For most of its life, Black Forest Labs have been known for still pictures. Its founders, Robin Rombach, Andreas Blattmann, and Patrick Esser, had helped build the latent diffusion work behind Stable Diffusion, then in 2024 opened a lab in Germany whose FLUX models became a default answer whenever someone wanted an image generator they could actually run.
This time, the lab changed the pitch.
In July it showed FLUX 3, a single system trained across images, video, audio, and the prediction of robot actions, and said the still-image piece would follow in the coming weeks.
And that product finally arrived to the masses.
The company introduced 'FLUX 3 Image' and described it as a way to place objects with bounding boxes, edit one region across several turns without touching the rest of the frame, generate at up to 4K, and compose a picture from as many as ten reference images.
Commercial weights are available now for those who want to fine-tune and host the model themselves. An open-weights version, the company said, is coming in the next few weeks.
The documentation is more specific than the thread, and a little less absolute.
Generation and editing share one endpoint. There is no mode switch. The prompt decides. Boxes are not a separate parameter. They are a JSON list appended to the prompt, each one written as top, left, bottom, right on a grid that runs from 0 to 1,000.
The same boxes are how the model is told to recolor, move, replace, or remove an element while leaving the unmarked area alone.
Reference images, between one and ten, can be passed as a URL or as base64, and the prompt is supposed to say what each one contributes.
Output sizes run from a 768-pixel square preview through roughly 1, 1.5, 2, and 4 megapixel classes, the last of those about 16 megapixels, across 15 aspect ratios. The product page shows a soba-shop scene rendered at 5,456 by 3,072 and treats native resolution, rather than an upscale after the fact, as the point.
List prices on the API are flat per image and do not change if references are attached: $0.041 at 768 square, $0.048 at 1K, $0.10 at 2K, and $0.607 at 4K.
During launch, the model page does not publish scores against other image systems, and the line about leaving every other pixel untouched is the company's description of the editing behavior, not a result from an independent test.
The release also closes a gap the lab opened on purpose.
When FLUX 3 debuted on July 23, video with native audio was the part in early access, action prediction was going out through partners, and image generation was listed as later. Video later gained 2K and 4K output.
On September 23, the lab open-sourced FLUX 3 Action, a 7-billion-parameter model it said topped NVIDIA's RoboLab benchmark. Image, the capability that made the company familiar in the first place, is the one that waited.
What shipped on Thursday is less a new aesthetic than a control surface.
Bounding boxes are a return to the kind of layout instruction graphic design already uses, a headline here, a face there, a panel in a grid, rather than a hope that a paragraph will arrange itself. Whether the model holds those boxes once the scene gets crowded is something the demos suggest and the docs do not measure.
For now the practical split is the one Black Forest Labs has used all year.
The API is live, the commercial license is a sales conversation, and the weights that earlier FLUX users downloaded and fine-tuned themselves are still a few weeks out.
None of those controls are new on their own. Local edits, reference images, and 4K output were already table stakes before this, and the published comparisons do not yet include this model.
On an edit that leaves the rest of the frame alone, OpenAI's GPT Image 2.5 is the system most often credited with doing it from a plain instruction. Google's Nano Banana models document up to 14 references and native 4K. ByteDance's Seedream 4.5 takes 10 references and also renders at 4K. Midjourney's V8 edit model stops at four.
Ten references puts FLUX 3 Image in the middle of that pack, and its $0.607 list price for 4K is the expensive end.
This puts FLUX 3 Image closer to Ideogram 4.0, an open-weight model from June that also takes bounding boxes in the prompt and publishes a layout score. It tops out at 2K and is built for type and posters.























































































































































































































































































































































































