World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say Atlas unifies the two things computer vision has always kept apart:
Justin: "Historically, reconstruction has been its own subfield in computer vision with its own specialized tasks and models. Generation is what all the text-to-video models are really good at."
"Those are great for creative applications if I want to imagine something that's never been there before."
"But now with Atlas, for the first time, we're putting these two different parts of visual intelligence together in one model. It can do both 3D reconstruction and generation together in one architecture."
"We had to make it multimodal from the start. This thing natively works on text, images, videos, and camera poses as a native input to the model... It uses 3D as a native modality that it works on."
Fei-Fei: "Computer vision has been around for more than half a century... Our field traditionally has multiple tracks. You go to a computer vision conference: you have the pixel generation track, you have some recognition track, and you have a 3D reconstruction track."
"This is an elegant model that unifies the problem of reconstruction and generation by anchoring on viewpoints and viewpoint estimation. That's just incredibly powerful."