{"id":23671,"date":"2026-09-03T00:01:05","date_gmt":"2026-09-03T00:01:05","guid":{"rendered":"https:\/\/scannn.com\/atlas-a-world-model-for-spatial-intelligence\/"},"modified":"2026-09-03T00:01:05","modified_gmt":"2026-09-03T00:01:05","slug":"atlas-a-world-model-for-spatial-intelligence","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/atlas-a-world-model-for-spatial-intelligence\/","title":{"rendered":"Atlas: A World Model\u00a0for Spatial Intelligence"},"content":{"rendered":"\n<div>\n<!-- --><\/p>\n<p>World models generate, reconstruct, and simulate any possible world.<br \/>\nThey understand how worlds appear, behave, and evolve<br \/>\nso that we can render imagined worlds for creative users,<br \/>\nsimulate the real world in high fidelity, and help robots plan actions.<br \/>\nAt World Labs, we build these general purpose world models in pursuit of spatial intelligence.<\/p>\n<p>Today we are introducing Atlas, our next-generation world model.<br \/>\nAtlas is an omni model that we pretrained from scratch to natively operate on<br \/>\ntext, images, video, and 3D.<br \/>\nIt is a multimodal autoregressive diffusion transformer:<br \/>\nall inputs are combined into a shared spatial context.<br \/>\nAtlas uses that context to generate what comes next,<br \/>\nstaying consistent in 3D with everything it has seen and imagining what lies beyond it.<br \/>\nAtlas is built to scale: its performance improves with increased training compute,<br \/>\nand we expect this trend to hold as we continue scaling.<\/p>\n<p>Atlas can perform a broad range of tasks<br \/>\nspanning world generation, reconstruction, and simulation:<\/p>\n<p><!-- --><\/p>\n<ul>\n<li><a href=\"#camera-controlled-generation\"><strong>Camera-Controlled Generation<\/strong><\/a>:<br \/>\nAtlas generates images and videos from one or more images with pixel-perfect camera control,<br \/>\noutputting up to 1 minute of video at 1440p.<\/li>\n<li><a href=\"#spatial-reconstruction\"><strong>Spatial Reconstruction<\/strong><\/a>:<br \/>\nAtlas reconstructs real world scenes from one to dozens of input images.<br \/>\nIt generates both image frames from novel views and explicit 3D outputs,<br \/>\noutperforming state-of-the-art models specialized for 3D reconstruction.<\/li>\n<li><a href=\"#space-time-simulation\"><strong>Space-Time Simulation<\/strong><\/a>:<br \/>\nAtlas models space and time from input videos,<br \/>\nreframing videos for dramatic visual effects<br \/>\nand enabling <a href=\"https:\/\/www.worldlabs.ai\/blog\/.\/real-to-sim-to-real\">Real-to-Sim<\/a> workflows for robotics.<\/li>\n<li><a href=\"#image-generation\"><strong>Image Generation<\/strong><\/a>:<br \/>\nAtlas generates images and 360 panoramas from text;<br \/>\nit can follow complex prompts, render text,<br \/>\nand generate a wide variety of visual styles.<\/li>\n<\/ul>\n<p>Atlas will power future versions of <a target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\/\/marble.worldlabs.ai\/\">Marble<\/a><br \/>\nand other products from World Labs.<\/p>\n<h2 class=\"group\" id=\"camera-controlled-generation\"><span>Camera-Controlled Generation<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#camera-controlled-generation\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h2>\n<p>Atlas takes one or more reference images and generates<br \/>\nnew views at any camera position and angle you specify.<br \/>\nGenerated views match the content and geometry of the input images,<br \/>\nsmoothly extrapolating beyond them to imagine parts of the scene<br \/>\nnot visible in the inputs.<\/p>\n<p>Atlas handles a broad range of scene types, visual styles, and camera motions.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"full\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Videos are generated from<!-- --> <!-- --><br \/>\n<span class=\"text-indigo-500\">one to six input images<\/span><br \/>\nwith manually designed camera paths<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h3 class=\"group\" id=\"pixel-perfect-camera-control\"><span>Pixel-Perfect Camera Control<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#pixel-perfect-camera-control\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas uses precise camera geometry as a native input type,<br \/>\ngoing beyond coarse text-based instructions for camera control.<br \/>\nThis lets you frame every shot and control every motion.<\/p>\n<p>In the examples here,<br \/>\nAtlas generates a complete scene from <b>a single input image<\/b>.<br \/>\nIt uses the content of the input image along with its broad world knowledge<br \/>\nto imagine what the scene should look like from new angles.<br \/>\nFor example, it generates the back side of the robot,<br \/>\nand it guesses that there should be a grassy lawn next to the pool.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>From a single image, Atlas generates views from any angle. Drag to change<br \/>\nthe view.<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h3 class=\"group\" id=\"generating-with-spatial-context\"><span>Generating with Spatial Context<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#generating-with-spatial-context\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Similar to an LLM, Atlas first encodes its inputs into a context,<br \/>\nthen generates outputs conditioned on the context.<br \/>\nHowever, Atlas is unique because each image is grounded at a 3D position in space;<br \/>\nthis forms a <b>spatial context<\/b>.<\/p>\n<p>Managing this spatial context unlocks entirely new kinds of creative control.<br \/>\nFor example, you can place two unrelated reference images in the context<br \/>\nand position them in 3D space; Atlas then generates a world that smoothly interpolates between them.<\/p>\n<p>These examples demonstrate the model&#8217;s world knowledge and creativity;<br \/>\nit imagines doorways, hallways, nooks, and other transitions between<br \/>\notherwise unrelated image pairs.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Select left and right frames to populate the spatial context, and Atlas<br \/>\nstitches them together<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h3 class=\"group\" id=\"controllable-long-videos\"><span>Controllable Long Videos<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#controllable-long-videos\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas lets you generate long videos with precise control<br \/>\nby combining camera movement and spatial context management.<br \/>\nYou design every scene and every camera angle.<br \/>\nThis puts you in the director&#8217;s chair:<br \/>\nyou are staging the scene, not pulling the lever of a slot machine.<\/p>\n<p>In the example below, we generate a 1 minute video at 1440p resolution using a small number of<br \/>\nreference images.<br \/>\nWe hand-design a camera path through the scene,<br \/>\nand Atlas generates a coherent world.<br \/>\nThe rest of the videos on this page have been compressed to optimize page performance.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h2 class=\"group\" id=\"spatial-reconstruction\"><span>Spatial Reconstruction<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#spatial-reconstruction\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h2>\n<p>Atlas reconstructs real-world spaces from one or more input images.<br \/>\nIt does not require special capture equipment or hundreds of dense views to<br \/>\nfaithfully reconstruct objects and scenes.<br \/>\nWe believe Atlas is a major step forward toward solving<br \/>\nthe problem of novel view synthesis from sparse input images,<br \/>\na decades-old fundamental problem in 3D computer vision.<\/p>\n<h3 class=\"group\" id=\"reconstructing-from-multiple-images\"><span>Reconstructing from Multiple Images<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#reconstructing-from-multiple-images\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas can take a variable number of input views of a scene.<br \/>\nWhen parts of the world are not visible in the input views,<br \/>\nAtlas imagines a plausible way to fill in the gaps<br \/>\nby drawing from its rich world knowledge.<\/p>\n<p>But sometimes you do not want imagination;<br \/>\nyou might want an exact reconstruction of a real-world location.<br \/>\nPassing more input images gives Atlas more context:<br \/>\nthe more it sees, the less it imagines.<br \/>\nAtlas typically gives faithful reconstructions with as few as two or three images,<br \/>\noutperforming state-of-the-art results by models specially trained only for 3D reconstruction.<br \/>\nHowever, Atlas can also make use of over a hundred input images in its spatial context,<br \/>\nallowing for faithful recreation of real world environments.<\/p>\n<p>In the first example below,<br \/>\nAtlas generates an aerial view of the scene from just a single ground-level photo.<br \/>\nThe garden visible in the single input photo is accurately recreated in the model output,<br \/>\nbut the rest of the scene is imagined.<br \/>\nAfter adding a second real-world input image of the cottage next to the garden,<br \/>\nthe model&#8217;s output shows both the garden and the cottage,<br \/>\nbut it still imagines the house to the left.<br \/>\nAfter adding a third input image of the main house,<br \/>\nthe entire scene is accurately depicted.<\/p>\n<p>In the second example, we build up Stanford&#8217;s Main Quad piece by piece,<br \/>\nbeginning with the grassy main entrance and ending with the colorful mosaics decorating<br \/>\nthe facade of Memorial Church.<br \/>\nThough Atlas only receives two to twenty-five ground-level input images,<br \/>\nit can generate paths from aerial views flying far above the campus.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"full\" data-not-typeset=\"true\" data-ui=\"article-figure\"><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h3 class=\"group\" id=\"reconstructing-diverse-paths\"><span>Reconstructing Diverse Paths<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#reconstructing-diverse-paths\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas can generate many different trajectories through the same scene,<br \/>\ngiving new perspectives on the same input images.<br \/>\nNo matter how many times you change the camera path, the scene stays consistent.<\/p>\n<p>In the example below, we show that given a small set of input images,<br \/>\nAtlas can generate multiple camera paths through the same scene.<br \/>\nDifferent camera paths can emphasize various parts of the scene<br \/>\nor change moods by varying in speed, length, or complexity.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Atlas can generate many different paths through the same scene using a<br \/>\nsmall set of input images.<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h3 class=\"group\" id=\"explicit-3d-outputs\"><span>Explicit 3D Outputs<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#explicit-3d-outputs\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>In the results above you have seen Atlas output 2D images and videos,<br \/>\nwhich are sufficient for some applications.<br \/>\nBut workflows in robotics, gaming, design, VFX, and beyond often require explicit 3D outputs.<br \/>\nAtlas natively operates on both 2D image frames and 3D depth maps,<br \/>\nenabling it to output worlds as point clouds or 3D Gaussian splats.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>From one image, Atlas generates new views and 3D geometry, then converts to<br \/>\n3D Gaussian splats<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<p>From a single input image,<br \/>\nAtlas produces a full 3D world by jointly generating new views and estimating their geometry.<br \/>\nFrom a video of a real space,<br \/>\nit predicts the depth of every frame and combines them into a 3D reconstruction.<br \/>\nIn either case, Atlas fills in regions that no camera ever saw.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Atlas can reconstruct 3D point clouds from input videos<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<p>Point clouds estimate a scene&#8217;s geometry, but 3D Gaussian splats make it usable.<br \/>\nAtlas fills the remaining gaps and turns the point cloud into a complete splat scene<br \/>\nthat renders on-device at high resolution and frame rates.<br \/>\nThis is the same representation used in <a target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\/\/marble.worldlabs.ai\/\">Marble<\/a>,<br \/>\nenabling Atlas to integrate naturally with the rest of our products.<\/p>\n<p><!-- --><\/p>\n<h2 class=\"group\" id=\"space-time-simulation\"><span>Space-Time Simulation<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#space-time-simulation\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h2>\n<p>Atlas serves as a world simulator.<br \/>\nIt understands both the spatial structure of the world<br \/>\nand how the world evolves over time.<br \/>\nCombining its spatial and temporal abilities leads to new applications<br \/>\nfor VFX, robotics, and beyond.<\/p>\n<h3 class=\"group\" id=\"reframing-video\"><span>Reframing Video<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#reframing-video\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas turns a handful of ordinary cameras into a &#8220;bullet time&#8221; multiview capture studio.<br \/>\nWith footage from as few as three cameras,<br \/>\nAtlas can freeze time and reframe shots,<br \/>\nletting you view events from impossible angles.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Real-world videos can be reframed from new camera angles without an<br \/>\nexpensive capture studio<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<p>Notably, these shots did not require professional photographers or specialized equipment.<br \/>\nEach of them was filmed by a few engineers and researchers<br \/>\nwith ordinary cell phones on tripods and clamps that fit in a backpack.<br \/>\nAtlas reconstructs the scene from three to five camera views,<br \/>\nafter which you can reframe shots however you like.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Behind the scenes: the clips above were captured using just a few cell<br \/>\nphones and action cameras<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h3 class=\"group\" id=\"robotics-simulation\"><span>Robotics Simulation<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#robotics-simulation\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas opens up new ways to scale <a href=\"https:\/\/www.worldlabs.ai\/blog\/.\/real-to-sim-to-real\">Real-to-Sim<\/a><br \/>\nfor both navigation and manipulation.<\/p>\n<p>You have already seen Atlas reconstruct a space in explicit 3D from a few images.<br \/>\nFor robotics, reconstruction is only half the job:<br \/>\nas a simulated robot moves through space,<br \/>\nAtlas also generates the RGB and depth data its sensors would observe along the way.<br \/>\nThe world and the robot&#8217;s view of it come from the same model.<\/p>\n<p>In these examples,<br \/>\nwe captured two large environments with a cell phone video, using 24 frames each for reconstruction.<br \/>\nScanning spaces like these traditionally requires elaborate and expensive equipment.<br \/>\nWe then simulate different kinds of robots navigating different paths<br \/>\nand use Atlas to generate images from the perspective of the robot&#8217;s body-mounted cameras.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Atlas reconstructs spaces and aids in simulating robot navigation<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<p>Robotic manipulation goes a step further.<br \/>\nFrom a few casual recordings,<br \/>\nAtlas aids in building a simulation that also captures how objects move and interact.<br \/>\nOnce a task is simulated, you can vary it easily:<br \/>\nchange the objects, their positions, the robot&#8217;s motion, the lighting, and the background.<br \/>\nThe result is diverse training data and testing environments for robotics at scale.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\">\n<div class=\"relative overflow-hidden rounded-lg ring-1 ring-black\/10\" data-figure-surface=\"true\">\n<div class=\"relative size-full\"><video aria-label=\"Robot manipulation demo showing sparse real-world capture, an Atlas simulation, and environment variations\" autoplay=\"\" class=\"block w-full\" loop=\"\" muted=\"\" playsinline=\"\" poster=\"https:\/\/wlt-ai-cdn.art\/atlas\/assets-827\/robotics\/poster\/atlas-robotics-demo-v3.jpg\" preload=\"metadata\" src=\"https:\/\/wlt-ai-cdn.art\/atlas\/assets-0831\/robotics\/optimized\/atlas-robotics-demo-v3.webm\"><\/p>\n<p>Browser does not support video playback.<\/p>\n<p><\/video><button type=\"button\" aria-label=\"Show video controls\" class=\"absolute inset-0 z-10 touch-manipulation bg-transparent outline-none focus-visible:ring-2 focus-visible:ring-white focus-visible:ring-inset\"\/><\/div>\n<\/div><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Atlas enables Real-to-Sim from just a few real-world recordings, recreating<br \/>\nphysical interactions with rigid, articulated, and deformable objects while<br \/>\nsupporting controllable variations.<\/p>\n<\/figcaption><\/figure>\n<h2 class=\"group\" id=\"image-generation\"><span>Image Generation<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#image-generation\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h2>\n<p>The primary focus of Atlas is world modeling,<br \/>\nand every image is a window to a possible world.<br \/>\nThough image generation is not its primary focus,<br \/>\nAtlas is a capable image generator:<br \/>\nit follows complex prompts, renders text, and generates a wide variety of visual styles.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<p>Atlas also generates 360 images from text or image prompts,<br \/>\nwhere again it can generate a wide variety of scene types and visual styles.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<h2 class=\"group\" id=\"technical-details\"><span>Technical Details<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#technical-details\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h2>\n<h3 class=\"group\" id=\"model-architecture\"><span>Model Architecture<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#model-architecture\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas is an omni model designed to handle many tasks and many kinds of input and output data in a<br \/>\nsingle unified architecture, putting spatial control at the heart of the model.<br \/>\nThese goals require us to depart from standard architectures used by both LLMs and<br \/>\nvideo models, and design a new base architecture to serve as the foundation of future world models.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-bleed=\"wide\" data-not-typeset=\"true\" data-ui=\"article-figure\"><figcaption class=\"text-grey-150 text-center text-sm font-medium text-balance first:mb-2 last:mt-2 lg:text-base [&amp;&gt;p]:m-0 [&amp;&gt;p]:leading-[inherit] mx-auto max-w-[45rem]\" data-slot=\"caption\">\n<p>Atlas is a multimodal autoregressive diffusion transformer. Its inputs are<br \/>\ngrounded in 3D space to form a spatial context, and it generates multimodal<br \/>\noutputs conditioned on its context.<\/p>\n<\/figcaption><!--$!--><template data-dgst=\"BAILOUT_TO_CLIENT_SIDE_RENDERING\"\/><!--\/$--><\/figure>\n<p>Atlas is a <strong>multimodal autoregressive diffusion transformer<\/strong>.<br \/>\nIt operates on multimodal sequences, generating each new element of the sequence one at a time.<br \/>\nThese architectural properties work together to achieve our goals,<br \/>\nand taken together they enable a new paradigm of generation based on a <strong>spatial context<\/strong>.<br \/>\nWe unpack these ideas in turn:<\/p>\n<ul>\n<li><strong>Multimodal<\/strong>: Atlas can natively process many different data types.<br \/>\nAt present it can operate on text, images, camera poses, and 3D depth maps;<br \/>\nvideos are represented as sequences of images.<br \/>\nEach image and depth map is conditioned on an explicit camera pose,<br \/>\nmaking spatial control a central component of the architecture.<\/li>\n<li><strong>Autoregressive<\/strong>: Atlas operates on sequences of elements,<br \/>\nwhere each element is one of the multimodal data types above.<br \/>\nEach output is generated one at a time, conditioned on earlier parts<br \/>\nof the sequence.<br \/>\nThis flexible design naturally adapts to a wide variety of tasks:<br \/>\neach task is just a different kind of sequence, where inputs are followed by outputs.<\/li>\n<li><strong>Diffusion<\/strong>: Atlas is a rectified flow model that generates outputs<br \/>\nby gradually denoising them.<br \/>\nDiffusion models excel at modeling high-dimensional continuous data like<br \/>\nimages and video,<br \/>\nand can naturally trade off speed and quality by varying the number of denoising steps<br \/>\nused during inference.<\/li>\n<li><strong>Transformer<\/strong>: The transformer architecture consists primarily of large matrix<br \/>\nmultiply operations and is well-adapted to modern hardware.<br \/>\nIt is a robust backbone for world modeling.<\/li>\n<\/ul>\n<p>Atlas is a blend of ideas from modern LLMs and video models.<br \/>\nIt can benefit from architectural, algorithmic, and systems advances<br \/>\nused in both types of models.<\/p>\n<p>Like an LLM, it is an autoregressive transformer,<br \/>\nso it can take advantage of innovations used to serve and accelerate LLMs<br \/>\nincluding KV-caching, cache-aware routing, disaggregated serving, and more.<br \/>\nLike a modern image or video model, it is a latent diffusion model<br \/>\nand can make use of algorithms such as diffusion distillation, classifier-free guidance,<br \/>\nshifted noise schedules, and advances in VAE design.<\/p>\n<h3 class=\"group\" id=\"benchmarks\"><span>Benchmarks<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#benchmarks\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Atlas is an omni model for world modeling that performs many tasks.<br \/>\nThere is thus no single benchmark that fully captures its generality.<br \/>\nWe highlight quantitative evaluations of Atlas on two key tasks:<br \/>\ncamera-conditioned generation and 3D reconstruction.<br \/>\nOn both tasks it outperforms more specialized models.<\/p>\n<p>We compare against a selection of top-performing video models for camera-conditioned generation.<br \/>\nIn each trial, we pair a single input image with a sequence of one to three cinematic camera motions<br \/>\n(pan, truck, crane, etc.).<\/p>\n<p>We prompt each model with a single input image and a target camera path.<br \/>\nFor Atlas, we encode the camera path using its native camera input format.<br \/>\nOther models do not accept cameras as a native input format,<br \/>\nso we describe the camera path in the input text prompt,<br \/>\nusing standard cinematic terms.<br \/>\nIt is possible that more sophisticated prompt engineering or creative multimodal prompts<br \/>\ncould improve camera following for some models,<br \/>\nbut we use text as it is the most common input modality for describing camera motions.<\/p>\n<p>Third-party human raters judge which model better follows the intended camera path.<br \/>\nThese results confirm that <b>Atlas outperforms recent video models at camera-controlled generation<\/b>,<br \/>\nand this advantage grows as camera trajectories become more complex.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-not-typeset=\"true\" data-ui=\"article-figure\">\n<div class=\"flex w-full flex-col gap-4 \" data-ui=\"benchmark-chart\">\n<div class=\"flex items-center justify-between gap-4\">\n<p class=\"text-foreground m-0 font-sans text-base leading-normal font-semibold\">Camera-Controlled Generation<\/p>\n<p><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" aria-hidden=\"true\" class=\"text-foreground h-5 w-auto shrink-0\" fill=\"currentColor\" viewbox=\"0 0 276 225\"><path d=\"M.251734 8.97818 35.5517 182.978c1.5 7.5 9.9 9.1 15 3.4l15.2-15.4c13.8-14 34.2-12.6 46.7003 2.5l48.1 49.2c3.5 3.7 10.2 2.3 12.6-2.2l75.2-133.9998c3.4-6-1.7-11.7-8.4-9.9l-120.8 42.7998c-11.3 4.8-24.2003 10.1-38.9003-6.5l-66.8-109.89982c-4.79997-6.3-14.79996-1.8-13.299964 6h.099998Z\"\/><\/svg><\/div>\n<p><svg aria-label=\"Share of voters choosing Atlas over each model: MiniMax H3 75 percent, with a 95 percent confidence interval from 68 to 82; Gemini Omni Flash 81 percent, with a 95 percent confidence interval from 74 to 87; Happy Horse 1.1 86 percent, with a 95 percent confidence interval from 82 to 91; FLUX 3 93 percent, with a 95 percent confidence interval from 88 to 97; Seedance 2.5 94 percent, with a 95 percent confidence interval from 90 to 97. 50 percent means the vote was evenly split.\" class=\"block w-full font-sans\" height=\"256\" role=\"img\" viewbox=\"0 0 720 256\" width=\"720\"><g class=\"text-grey-150\" fill=\"currentColor\" font-size=\"14\"><text text-anchor=\"start\" x=\"140\" y=\"14\">\u2190 Other model preferred<\/text><text text-anchor=\"end\" x=\"702\" y=\"14\">Atlas preferred \u2192<\/text><\/g><g><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" text-anchor=\"end\" x=\"128\" y=\"53\">MiniMax H3<\/text><g class=\"text-foreground\" stroke=\"currentColor\" stroke-opacity=\"0.5\"><line x1=\"365.1746666666668\" x2=\"509.1965333333334\" y1=\"48\" y2=\"48\"\/><line x1=\"365.1746666666668\" x2=\"365.1746666666668\" y1=\"43\" y2=\"53\"\/><line x1=\"509.1965333333334\" x2=\"509.1965333333334\" y1=\"43\" y2=\"53\"\/><\/g><circle cx=\"439.5085333333334\" cy=\"48\" fill=\"var(--indigo-500)\" r=\"5.5\"\/><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" font-weight=\"700\" text-anchor=\"middle\" x=\"439.5085333333334\" y=\"37\">75<!-- -->%<\/text><\/g><g><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" text-anchor=\"end\" x=\"128\" y=\"89\">Gemini Omni Flash<\/text><g class=\"text-foreground\" stroke=\"currentColor\" stroke-opacity=\"0.5\"><line x1=\"432.53973333333334\" x2=\"567.2698666666666\" y1=\"84\" y2=\"84\"\/><line x1=\"432.53973333333334\" x2=\"432.53973333333334\" y1=\"79\" y2=\"89\"\/><line x1=\"567.2698666666666\" x2=\"567.2698666666666\" y1=\"79\" y2=\"89\"\/><\/g><circle cx=\"502.2277333333332\" cy=\"84\" fill=\"var(--indigo-500)\" r=\"5.5\"\/><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" font-weight=\"700\" text-anchor=\"middle\" x=\"502.2277333333332\" y=\"73\">81<!-- -->%<\/text><\/g><g><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" text-anchor=\"end\" x=\"128\" y=\"125\">Happy Horse 1.1<\/text><g class=\"text-foreground\" stroke=\"currentColor\" stroke-opacity=\"0.5\"><line x1=\"509.1965333333332\" x2=\"606.7597333333332\" y1=\"120\" y2=\"120\"\/><line x1=\"509.1965333333332\" x2=\"509.1965333333332\" y1=\"115\" y2=\"125\"\/><line x1=\"606.7597333333332\" x2=\"606.7597333333332\" y1=\"115\" y2=\"125\"\/><\/g><circle cx=\"560.3010666666665\" cy=\"120\" fill=\"var(--indigo-500)\" r=\"5.5\"\/><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" font-weight=\"700\" text-anchor=\"middle\" x=\"560.3010666666665\" y=\"109\">86<!-- -->%<\/text><\/g><g><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" text-anchor=\"end\" x=\"128\" y=\"161\">FLUX 3<\/text><g class=\"text-foreground\" stroke=\"currentColor\" stroke-opacity=\"0.5\"><line x1=\"581.2074666666666\" x2=\"669.4789333333333\" y1=\"156\" y2=\"156\"\/><line x1=\"581.2074666666666\" x2=\"581.2074666666666\" y1=\"151\" y2=\"161\"\/><line x1=\"669.4789333333333\" x2=\"669.4789333333333\" y1=\"151\" y2=\"161\"\/><\/g><circle cx=\"629.9890666666669\" cy=\"156\" fill=\"var(--indigo-500)\" r=\"5.5\"\/><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" font-weight=\"700\" text-anchor=\"middle\" x=\"629.9890666666669\" y=\"145\">93<!-- -->%<\/text><\/g><g><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" text-anchor=\"end\" x=\"128\" y=\"197\">Seedance 2.5<\/text><g class=\"text-foreground\" stroke=\"currentColor\" stroke-opacity=\"0.5\"><line x1=\"597.4680000000001\" x2=\"674.1247999999999\" y1=\"192\" y2=\"192\"\/><line x1=\"597.4680000000001\" x2=\"597.4680000000001\" y1=\"187\" y2=\"197\"\/><line x1=\"674.1247999999999\" x2=\"674.1247999999999\" y1=\"187\" y2=\"197\"\/><\/g><circle cx=\"639.2808\" cy=\"192\" fill=\"var(--indigo-500)\" r=\"5.5\"\/><text class=\"text-foreground\" fill=\"currentColor\" font-size=\"14\" font-weight=\"700\" text-anchor=\"middle\" x=\"639.2808\" y=\"181\">94<!-- -->%<\/text><\/g><g class=\"text-grey-150\" stroke=\"currentColor\"><line stroke-opacity=\"0.35\" x1=\"140\" x2=\"155.67000000000002\" y1=\"212\" y2=\"212\"\/><line stroke-opacity=\"0.35\" x1=\"163.67000000000002\" x2=\"702\" y1=\"212\" y2=\"212\"\/><line stroke-opacity=\"0.7\" x1=\"153.67000000000002\" x2=\"157.67000000000002\" y1=\"216\" y2=\"208\"\/><line stroke-opacity=\"0.7\" x1=\"161.67000000000002\" x2=\"165.67000000000002\" y1=\"216\" y2=\"208\"\/><\/g><g class=\"text-grey-150\" fill=\"currentColor\" font-size=\"14\"><text text-anchor=\"start\" x=\"140\" y=\"228\">0<\/text><text text-anchor=\"middle\" x=\"179.34\" y=\"228\">50<\/text><text text-anchor=\"middle\" x=\"283.872\" y=\"228\">60<\/text><text text-anchor=\"middle\" x=\"388.404\" y=\"228\">70<\/text><text text-anchor=\"middle\" x=\"492.9359999999999\" y=\"228\">80<\/text><text text-anchor=\"middle\" x=\"597.4680000000001\" y=\"228\">90<\/text><text text-anchor=\"end\" x=\"702\" y=\"228\">100<\/text><text font-size=\"14\" text-anchor=\"middle\" x=\"421\" y=\"250\">Share of voters choosing Atlas<\/text><\/g><\/svg><\/p>\n<\/div>\n<\/figure>\n<p>We additionally evaluate Atlas on the task of 3D reconstruction from sparse input views.<br \/>\nIn each trial, the model receives a set of images and their camera poses,<br \/>\nand predicts a 3D point corresponding to each input pixel.<br \/>\nThis problem has attracted much interest in the academic community,<br \/>\nand many specialist reconstruction models have been developed in recent years.<\/p>\n<p>Atlas is an omni model which performs both generation and reconstruction.<br \/>\nDespite its generality, <b>Atlas outperforms the best specialized open-source reconstruction models<\/b>.<\/p>\n<p>We evaluate on several state-of-the-art benchmarks for this task,<br \/>\nreproducing the results for all baselines<br \/>\nto ensure a common and fair evaluation protocol across all methods.<\/p>\n<figure class=\"text-foreground my-8 w-full text-base leading-normal [&amp;+h2]:mt-6 [&amp;+h3]:mt-4 [&amp;+p]:mt-2 sm:my-12 \" data-not-typeset=\"true\" data-ui=\"article-figure\">\n<div class=\"flex flex-col gap-8\">\n<div class=\"flex items-center justify-between gap-4\">\n<p class=\"text-foreground m-0 font-sans text-base leading-normal font-semibold\">3D Reconstruction Error (lower is better)<\/p>\n<p><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" aria-hidden=\"true\" class=\"text-foreground h-5 w-auto shrink-0\" fill=\"currentColor\" viewbox=\"0 0 276 225\"><path d=\"M.251734 8.97818 35.5517 182.978c1.5 7.5 9.9 9.1 15 3.4l15.2-15.4c13.8-14 34.2-12.6 46.7003 2.5l48.1 49.2c3.5 3.7 10.2 2.3 12.6-2.2l75.2-133.9998c3.4-6-1.7-11.7-8.4-9.9l-120.8 42.7998c-11.3 4.8-24.2003 10.1-38.9003-6.5l-66.8-109.89982c-4.79997-6.3-14.79996-1.8-13.299964 6h.099998Z\"\/><\/svg><\/div>\n<div class=\"flex flex-col gap-4\">\n<div aria-hidden=\"true\" class=\"flex flex-col gap-1.5\">\n<p><span class=\"text-grey-150 flex items-center gap-2 font-sans text-sm font-medium\"><span class=\"size-2.5 rounded-[2px]\" style=\"background:var(--indigo-500)\"\/><span>Atlas<!-- --> <strong class=\"text-foreground font-semibold\">(Ours)<\/strong><\/span><\/span><span class=\"text-grey-150 flex items-center gap-2 font-sans text-sm font-medium\"><span class=\"size-2.5 rounded-[2px]\" style=\"background:oklch(0.84 0.07 340)\"\/><span>Pi3X (posed)<\/span><\/span><span class=\"text-grey-150 flex items-center gap-2 font-sans text-sm font-medium\"><span class=\"size-2.5 rounded-[2px]\" style=\"background:oklch(0.8 0.038 175)\"\/><span>\u03c0\u00b3<\/span><\/span><span class=\"text-grey-150 flex items-center gap-2 font-sans text-sm font-medium\"><span class=\"size-2.5 rounded-[2px]\" style=\"background:oklch(0.78 0.04 250)\"\/><span>VGGT-\u03a9 1B<\/span><\/span><span class=\"text-grey-150 flex items-center gap-2 font-sans text-sm font-medium\"><span class=\"size-2.5 rounded-[2px]\" style=\"background:oklch(0.84 0.04 85)\"\/><span>Depth Anything 3<\/span><\/span><span class=\"text-grey-150 flex items-center gap-2 font-sans text-sm font-medium\"><span class=\"size-2.5 rounded-[2px]\" style=\"background:oklch(0.82 0.035 25)\"\/><span>MapAnything<\/span><\/span><\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/figure>\n<h3 class=\"group\" id=\"model-scaling\"><span>Model Scaling<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#model-scaling\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h3>\n<p>Most progress in modern AI has been driven by scaling.<br \/>\nModels improve in large part by scaling up simple algorithms to make use of more data and compute.<\/p>\n<p>We see strong evidence that Atlas will continue to improve with scale.<br \/>\nWe pretrained Atlas from scratch on a large diverse corpus of multimodal data.<br \/>\nOver the course of development,<br \/>\nwe trained a series of models of increasing size and training compute,<br \/>\nand found that each new level of compute unlocked new model capabilities.<br \/>\nWe are confident that our future world models will follow this trend,<br \/>\ndramatically improving their capabilities as we continue to scale.<\/p>\n<h2 class=\"group\" id=\"build-with-atlas\"><span>Build with Atlas<\/span><a aria-label=\"Copy link to this section\" class=\"text-grey-150 hover:text-foreground focus-visible:outline-foreground ease ml-2 inline-flex touch-manipulation rounded-sm align-[-0.125em] opacity-0 transition-[color,opacity] duration-200 outline-none group-hover:opacity-60 hover:opacity-100 focus-visible:opacity-100 focus-visible:outline-2 focus-visible:outline-offset-2 [@media(hover:none)]:opacity-60\" title=\"Copy link to this section\" href=\"#build-with-atlas\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"24\" height=\"24\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"1.75\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2 size-4.5\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/a><\/h2>\n<p>Atlas is entering early access with select partners.<br \/>\nIf you would like to build with it, request access below and we will reach out.<br \/>\nWe are excited to see what you build, and to work with you to make Atlas<br \/>\nthe go-to world model for generating, reconstructing, and simulating any world.<\/p>\n<p>We are also <a target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\/\/www.worldlabs.ai\/careers\">hiring<\/a> across research and engineering to advance spatial intelligence.<\/p>\n<p><em class=\"[font-synthesis:style]\">This post was produced by the World Labs team. Please cite as:<\/em><\/p>\n<div class=\"relative\">\n<pre class=\"bg-foreground\/[0.05] text-foreground m-0 overflow-x-auto rounded-md px-5 py-6 font-mono text-xs leading-5 whitespace-pre-wrap sm:px-7 sm:text-[0.8125rem]\"><code>@article{worldlabs2026atlas,\n    author = {World Labs Team},\n    title = {Atlas: A World Model for Spatial Intelligence},\n    journal = {World Labs Blog},\n    year = {2026},\n    note = {https:\/\/www.worldlabs.ai\/blog\/atlas},\n}<\/code><\/pre>\n<p><button aria-label=\"Copy citation to clipboard\" class=\"absolute top-2.5 right-2.5 flex size-8 cursor-pointer items-center justify-center rounded-md bg-background\/80 ring-foreground\/15 text-grey-150 shadow-sm ring-1 backdrop-blur-sm hover:text-foreground hover:bg-background transition-colors duration-150 ease-out outline-none focus-visible:outline-2 focus-visible:outline-offset-2 focus-visible:outline-foreground\" type=\"button\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" aria-hidden=\"true\" fill=\"none\" viewbox=\"0 0 24 24\" class=\"size-4\"><path d=\"M7.75 7.75V6.75C7.75 5.09315 9.09315 3.75 10.75 3.75H17.25C18.9069 3.75 20.25 5.09315 20.25 6.75V13.26C20.25 14.9169 18.9069 16.26 17.25 16.26H16.25M3.75 10.75V17.25C3.75 18.9069 5.09315 20.25 6.75 20.25H13.25C14.9069 20.25 16.25 18.9069 16.25 17.25V10.75C16.25 9.09315 14.9069 7.75 13.25 7.75H6.75C5.09315 7.75 3.75 9.09315 3.75 10.75Z\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"1.5\"\/><\/svg><\/button><span aria-live=\"polite\" class=\"sr-only\"\/><\/div>\n<\/div>\n<p><a href=\"https:\/\/www.worldlabs.ai\/blog\/atlas?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>World models generate, reconstruct, and simulate any possible world. They understand how worlds appear, behave, and evolve so that we can render imagined worlds for creative users, simulate the real world in high fidelity, and help robots plan actions. At World Labs, we build these general purpose world models in pursuit of spatial intelligence. Today [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23672,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23671","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23671","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23671"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23671\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23672"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23671"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23671"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23671"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}