Black Forest Labs spent its first two years being treated as an image company.
The Freiburg lab was founded in 2024 by researchers who had left Stability AI after working on Stable Diffusion, among them Robin Rombach, Andreas Blattmann and Patrick Esser.
FLUX.1 and later FLUX.2 made the name.
Video arrived only in July 2026 as part of FLUX 3, with generation going generally available in August. Then, Black Forest Labs released a minor update to FLUX 3, but with significance implication to how the product can be used.
Calling it the 'Flux Video Edit, the piece is smaller than that launch but more specific: it takes an MP4 file and a sentence.
To use this Flux Video Edit, users can give it a clip by URL or as base64, and then give it a prompt of up to 4,096 characters.
The model can then add, remove or replace objects and characters, swap a background, rewrite on screen text, change colors, materials or effects, change or translate spoken lines with lip movement adjusted to the new words, and alter the events of a shot.
Several of those changes can be stacked in one prompt.
After that, the system returns with another MP4 of the same length and aspect ratio, normalized to 24 frames per second, at up to 720p.
Official examples use a harbor clip in which a bucket is removed, an apron is recolored, crate lettering is rewritten, a seagull is added and snow is introduced.
Sources above 720p are downscaled. Clips longer than 15 seconds or larger than 50 MiB are rejected rather than trimmed. A 10 second edit is described as taking about 50 seconds.
Billing is $0.03 per second of output, so a 10 second clip costs $0.30.
Rejected jobs are not charged.
Source audio is kept unless the prompt asks for a change.
The documented edits are local rather than generative from scratch.
The lab's own claim is that competing editors often rewrite parts of the frame that were never mentioned. That claim sits on an internal Bradley Terry Elo chart published with the announcement, which places FLUX Video Edit [fast] near Google's Omni 1.1 and Alibaba's Wan 3 on quality at a lower cost per second.
There are some caveats, and some are bigger than others.
For example, a second video cannot be used as a style or motion reference. Then, still images and standalone audio are not accepted as conditioning. Masks are also not supported.
Duration, resolution, aspect ratio and frame rate cannot be set by the caller. New dialogue has to fit the length of the original spoken line. Stacking many edits in one pass can thin out fine detail.
Native output stops at 720p.
Higher resolution is left to FLUX Video Upscale, the August endpoint that can also take footage not made by FLUX.
Those gaps match what the company listed as not ready: video as reference, plus image and audio references.
That narrowness is the contrast with the larger video labs. Google’s Gemini Omni line, including Omni Flash, has been treated as a general audiovisual editor as much as a generator: natural language changes, character consistency across successive edits, and a high place on public text to video boards.
ByteDance’s Seedance 2.0 and 2.5 are built around longer single passes, up to about 30 seconds on 2.5, and dense reference boards that can take many images, clips and audio files, plus timestamp level control.
MiniMax’s H3 family has shown up near the top of video editing arenas, with H3 Max often discussed as a heavier, more controllable edit model than the base H3.
Alibaba’s Wan 3 takes mixed inputs, including documents in some workflows, and appears on Black Forest Labs' own cost quality chart as a higher spend option. Runway’s Aleph and Gen line still sit in the professional edit conversation as restyle and object tools with an established production audience.
OpenAI’s Sora 2 and Kuaishou's Kling 3 are more often scored on generation, physics and multi shot control than on cheap local edits, though both can rewrite a clip when asked.
In the end, FLUX Video Edit does not try to match those systems on duration, native 4K, or reference count. Fifteen seconds, 720p output and no reference media put it below Seedance, Wan and Omni on input flexibility.
The pitch is the other axis: a single prompt, no mask, $0.03 per second, and an internal chart that says the unused parts of the frame hold still.
t.What shipped is an API that takes a short MP4 and a sentence, returns another short MP4, and charges by the second.
Google, ByteDance, MiniMax and Alibaba already sell versions of that sentence.
Black Forest Labs is selling a cheaper, shorter, less conditioned one, and asking to be judged on how little else moves.






















































































































































































































































































































































































