Large language models proven themselves capable of learning how to write, to then paint still frames, to then invent short clips of motion.
What they still do poorly is treat a room, a courtyard, or a factory floor as a place rather than a picture. A new camera angle can collapse geometry. A robot that needs to train on a space it has never entered still depends on expensive scans or hand-built simulators.
That practical gap is what researchers now group under spatial intelligence and world models: systems that keep a consistent account of objects, viewpoints, and time so that generation, reconstruction, and simulation are not three separate crafts.
And World Labs, the San Francisco company co-founded by Fei-Fei Li, Justin Johnson, Ben Mildenhall, and Christoph Lassner, has been working in that category since it left stealth in 2024.
Its first public product, Marble, arrived in 2025 as a way to turn text, photos, video, or panoramas into persistent, exportable 3D environments. A research preview called RTFM later showed real-time frame generation without an explicit geometric store. In February 2026 the company raised an additional one billion dollars from investors that included AMD, Autodesk, Nvidia, Fidelity, and Sea. In July it acquired SceniX, a move framed around robotics simulation.
Those steps set the context for the model it posted on 1 September.
The new system is called 'Atlas.'
World Labs describes it as an omni world model pretrained from scratch so that text, images, video, camera poses, and depth maps sit in one shared spatial context.
The architecture is a multimodal autoregressive diffusion transformer that uses a rectified flow process to denoise outputs.
Each image or depth map is pinned to an explicit camera pose.
Generation then proceeds as a sequence conditioned on that layout, rather than as a video model steered only by a written description of a camera move.
The company says this is how Atlas can follow a designed trajectory at up to 1440p for as long as a minute from one to six reference photographs.
In the public thread and blog, the demonstrations fall into a few families.
One is camera-controlled generation: a handful of stills plus a hand-specified path produce a continuous shot, including aerial views that were never photographed, as in a Stanford Main Quad example built from ground-level frames.
Another is reconstruction.
Atlas predicts a 3D point per pixel and can emit point clouds or Gaussian splats. The company reports that two or three images are often enough for a reconstruction it calls faithful, while a single image matches only what was visible and invents the rest.
Adding views reduces that invention. A third family is space and time.
Footage from three to five ordinary cameras can be reframed into orbits that no physical rig captured, a cheap stand-in for classic bullet-time setups.
The same pipeline is shown producing the RGB and depth a robot sensor would see on an arbitrary trajectory through a space reconstructed from casual phone video.
World Labs published its own comparisons.
In a human preference test of camera-path following, third-party raters chose Atlas over several video generators by large margins. Those rivals received text descriptions of the desired motion; Atlas received native camera geometry, so the comparison is not symmetric.
On sparse-view reconstruction, the company reran open-source baselines and reported a lower average absolute-relative point-map error than models built only for that task, including Pi3X, VGGT-Ω, Depth Anything 3, and MapAnything.
Independent replication has not appeared. Some dataset-level numbers in secondary write-ups are closer or mixed. No paper, model card, or training-data inventory accompanied the launch.
Most systems grouped as world models still do one job.
Video models such as Gemini Omni Flash, FLUX 3, MiniMax H3, and Seedance 2.5 generate motion from text and frames, so camera movement is specified in language and 3D structure is not exported. Interactive models such as DeepMind Genie 3 and NVIDIA Cosmos produce an action-conditioned video stream that an agent can step through in real time.
World Labs’ earlier product, Marble, generated a persistent scene and exported Gaussian splats or meshes for use in other software. Atlas is a single model pretrained on text, images, video, camera poses, and depth. Each image is placed at a pose in a shared spatial context.
From that context it can generate video along a specified camera path, predict per-pixel 3D and output point clouds or Gaussian splats, reframe multi-camera footage, or render the RGB and depth a robot camera would record on a new trajectory.
Company evaluations compared it with video models on camera-path following and with reconstruction systems such as Pi3, VGGT, Depth Anything 3, and MapAnything on sparse-view geometry.
Those video comparisons gave Atlas native poses and the other models text descriptions. The reconstruction scores were run by World Labs. Atlas is not a real-time interactive simulator of the Genie type, and it has not been independently benchmarked.
Atlas is not inside Marble yet.
Existing Marble worlds and exports are unchanged. The company says the new model will power later versions of that product and other tools, and that it is opening early access with unnamed partners. A waitlist form is live.
Fei-Fei Li, posting the same day, called it a single model for generation and reconstruction and listed VFX and robotics as near-term uses.
NVIDIA researcher Jim Fan, among others on the thread, pointed to the real-to-sim path for robots as the part that matters most if the reconstructions hold up outside curated demos.
What remains open is the usual gap between a controlled announcement and production use.
Unseen surfaces still have to be imagined. Long camera paths still have to stay consistent after the last reference pixel. Benchmarks run by the lab that built the model will be read as a starting claim, not a settled ranking.
For now Atlas is another data point in a crowded year of world-model work: one system that tries to keep pixels, poses, and geometry in the same context, and that asks whether a few photographs can stand in for a scan, a stage, or a training hall.





















































































































































































































































































































































































