Most of the time spent on a video is not the idea. It is the stretch after the first rough assembly, when a cut is already too long, a caption lands half a second late, the music stays loud over a sentence, and the only way to fix any of it is to open the project and touch the pieces one by one.
AI-powered generative tools shortened the path to a first image or a first shot.
They did less for the pile of decisions that turn that shot into something a person would actually post.
Manus is playing its card again, and that is by showing a video editor inside Manus Studio, the desktop app that shipped with Manus 2.0 in late September.
In the example Manus chose, the person behind the product came back from a trip with 125 clips and almost two hours of footage. One prompt produced a 12:48 first cut that followed the itinerary rather than a highlight reel of pretty frames.
The blog post walks through the same project in more detail.
The agent imported the local files, transcribed them, sampled frames, arranged a story, then shortened the cut to 10:50 by dropping more than twenty slow or repeated spots, ending at 195 video segments.
Captions, text effects, music cues, and sound effects stayed on separate tracks: 325 captions, 112 text effects and stickers, seven music cues, and 118 sound effects, 757 elements in the project data Manus published.
The point of keeping those elements separate is the part that differs from a chat box that returns an MP4.
A small change, a reworded caption or a half-second trim, does not require regenerating the whole film and hoping the parts that already worked survive. The user can drag and trim on the timeline, or ask the agent to do the same thing. Both paths write to the same project, so a later instruction can continue from the human edit instead of starting over.
Studio runs on macOS and Windows.
A web session can start a video from a prompt, then hand the user toward the desktop app for the timeline.
Manus framed the launch as a consequence of what a general agent can already do around the edit.
In a follow-up post it listed browsing for what is circulating, writing code for motion graphics, layouts, and map routes, and routing each step to a different model or tool.
The use cases it named are concrete rather than open-ended: turn one video into versions in eight languages, build a launch film from Figma files and code repositories, cut hours of footage into a finished piece, and produce a set of user-generated-style ads for a product.
The blog also describes the agent searching news and social posts, scripting, generating shots, and checking its own renders, with an Alchemy mode that treats the agent as the first creative director before the timeline is handed back.
These figures come from Manus's own demo and write-up, not from an outside test.
A travel film assembled by the team that built the tool is a useful stress case. It is not a measure of how the same workflow behaves on messy audio, unclear briefs, or footage the agent has never seen.























































































































































































































































































































































































