David Fattal on Stereo Data, World Models, 3D Displays, and the Future of Spatial AI.
Subscribe: Apple • Spotify • Overcast • Pocket Casts • YouTube • AntennaPod • Podcast Addict • Amazon • RSS.
Ben Lorica talks with David Fattal, founder and CTO of Leia Inc., about why stereo data could become an important ingredient for AI systems that need to understand the physical world. They explore the limits of monocular video and synthetic data, the connection between 3D displays and a new source of training data for world models and robotics, and why phones and laptops may eventually capture and display spatial content by default. Fattal also explains why he thinks today’s 3D technology has a better chance of succeeding than the failed 3D TV wave.
Interview highlights – key sections from the video version:
-
-
- Introduction to David Fattal, Leia, and Immersity
- Stereo vs. Monocular Data: Why Two Eyes Matter
- Why World Models May Need Stereo Data
- Would Robotics Teams Use More Stereo Data?
- Recovering 3D Geometry From Monocular Video
- The Limits of Synthetic Data for Spatial Intelligence
- Why Proprietary Data Is Becoming the AI Moat
- Immersity, 3D Screens, and the Device-Data Flywheel
- A Future Where User-Generated Content Is Naturally 3D
- Why True Immersion Requires New Display Hardware
- Why 3D TV Failed and Why This Time Could Be Different
- Consumer Platforms, Dedicated Headsets, and Where 3D Fits
- The One-, Two-, and Three-Year Roadmap for 3D Displays
- Self-Driving Cars: Cameras vs. Lidar and Sensor Fusion
- Minority Report and a Vision for the Spatial Computing Future
-
Related content:
- A video version of this conversation is available on our YouTube channel.
- “World Model” is a mess. Here’s how to make sense of it.
- Foundation Models in Robotics: From Bespoke Machines to Generalist Brains
- Yoon Kim → Why Video Is AI’s Next Great Frontier
- Changan Chen → From Web Video to Real-World Robots
- Jeff Hawke → World Models Are Here—But It’s Still the GPT-2 Phase
- Ameet Talwalkar → Why Observability May Be AI’s Next Frontier
Support our work by subscribing to our newsletter📩
Transcript
Below is a polished and edited transcript.
Ben Lorica. All right. Today we have David Fattal, founder and CTO at Leia Inc. They have an offering called Immersity, which you can find at Immersity AI. The taglines are “Immersive experiences on any device” and “There is a new way to experience content, and it is immersive.” With Immersity, you can experience movies, images, and social media as if you’re inside the scene, right from your phone, tablet, monitor, or laptop. And with that, David, welcome to the podcast.
David Fattal. Hey, Ben. Thanks for having me.
Ben Lorica. All right. It sounds very promising. So I don’t have to buy one of these expensive headsets from Apple, huh? We’ll dig deeper into what Immersity does, but maybe I’ll start with the basics.
Obviously, this is a podcast with an audience primarily interested in AI, and as the audience will find out, there’s a bit of an overlap. In fact, what Immersity is doing seems to be central to a class of foundation models.
David, I think the audience has largely heard of what people call world models. World models have become somewhat of a marketing term, in that depending on who you talk to, they mean different things. There’s the world model where the model is better at understanding 3D spaces and how objects are laid out. Then there are models that specialize in prediction or evolution. What’s going to happen in this space if you drop this apple? And then there are models that go one step further and enable you to take action.
Based on my understanding of what you folks do, you aren’t necessarily building foundation models. You’re helping the foundation models by giving them better data.
David Fattal. Exactly. Better data, and helping them express themselves and visualize their conception of space in a human-consumable form, so that we humans can understand what the model actually understands about 3D space.
Ben Lorica. One of the things I’ve come across in preparing for this podcast is this notion of stereo data. There’s monocular data, which I imagine most of the videos on YouTube are. I’m not sure exactly what a modern iPhone camera video is, but for our audience, maybe let’s start there. What are the basics of stereo versus monocular, and why is stereo important?
David Fattal. Yeah. I think 99.9% of the content that we know is monocular, meaning it has one image. When you look at any content on your phone today, your computer, your TV, your laptop, it’s all flat. We’ve been used to consuming content in a flat manner.
Our brain has evolved, in a sense, to be able to figure out space projected somehow onto a 2D screen. Stereo is not new. We have two eyes. When we experience the world, we actually see the world from two slightly different points of view, and our brain does a lot of the intuitive work of merging these two pictures together and giving us the impression of volume and space. We do all that intuitively.
So the concept of using stereo shouldn’t come as a surprise. In a sense, it’s the latest evolution of visual media. If you want to take us 150 years back, we started with black-and-white pictures. Then we had color pictures. Then we added movement, so we had moving pictures and black-and-white movies. Then we added sound. Then we added color TV. Then we added touch with the UI.
Effectively, the evolution of media has been about trying to get to a point where we can faithfully describe our experience of the real world. Our experience of the real world turns out to be through two eyes.
Stereo is essentially the ability to record content, or at the opposite end render content from a screen, in a way that respects that binocular, two-eye experience. You’re able either to record or project two different images into the viewer’s eyes so that you can experience the content in the same way that you experience the world, which means in a fully immersive, spatial way.
Ben Lorica. So the people building world models, foundation models, frontier AI models, do they want stereo data? Do they have the capability to absorb stereo data into their pipelines, models, or model architectures?
David Fattal. That’s extremely interesting. I don’t think there’s a consensus right now. As you said, first of all, there are several flavors of world models.
I think everybody agrees that a world model involves the ability of an AI to make some kind of representation of the world or of a scene that is given, and to predict and plan what’s going to happen next, so that you can start to plan actions for a robot or an AI, or even a human wearing a pair of glasses or looking at an immersive display.
You need to be able to train the AI to understand space and geometry. The very basic thing is this: take a single picture of a person and ask how far the person is from the camera. This is what we call an ill-posed problem, because I could essentially have a big head and be located far away from the camera, or I could have a small head, which is my case, and be very close to the camera.
This geometry problem is ill-posed when you have a single point of view. If you have a second point of view, and this is why humans evolved to have two eyes, you get more information. It’s very costly to have two eyes. They’re a very fragile part of the head. They constantly get irritated and so on. The reason evolution has carried us along with two eyes is because they’re extremely useful for many things.
If you want to thread a needle, perceive a threat, anticipate something, or catch an object, it’s much easier to do it from two points of view, when you have this immediate reconstruction of the accurate geometry of the scene, rather than having to guess from other cues and experience.
When you build these world models, there’s one class of people saying, “Hey, just by looking at a lot of 2D pictures and videos, and maybe observing people moving around, it’s enough to give the AI the ability to guess all of these cues.”
Then there’s the other camp, which I belong to, saying that all these videos and pictures are good. They’re plentiful, and that’s the advantage. They’re good for pre-training. They’re good for getting a model to understand textures, shading, shadows, movement, and so on.
But I think you want to ground them in reality. You want to somehow steer these models so they don’t hallucinate and produce fantasies. For that, I think you need real-world, accurate geometry. The easiest way is through stereo data.
Other devices can be useful. You could have lidar, which gives you a single picture with a depth map. That will also give you geometry, but it’s limited. You have it mounted on cars, which is why a lot of self-driving car model training uses lidar. But you’re not going to get your car inside a building or in the middle of the jungle and so on.
The next best thing is stereo, and it’s actually much denser. You have a lot more pixels in a stereo picture than you have in lidar. You get really accurate correspondence between objects. If you’re able to train, fine-tune, or post-train this network with real-world stereo data, the contention is that it’s going to ground the model in reality.
That’s what’s going to make the difference between an okay model that occasionally hallucinates and a model that is always constrained to act and understand reality as it is.
Ben Lorica. Wear your objective hat for a second here. Would you say that right now that’s a minority perspective? I’ve talked to a bunch of people building foundation models for robotics. Some of them call them world models, but the main application is robotics.
David Fattal. Yeah.
Ben Lorica. Most of them just use video. Mainly, they use virtual training where you have an operator mimicking a movement in a virtual world, but most of them use video. It could be because the type of data you just described, stereo data, is not plentiful.
David Fattal. It’s very rare.
Ben Lorica. Very rare. Now, if I go back to the same people and tell them, “Hey, here’s an unlimited amount of stereo data for the use case that you’re interested in…”
David Fattal. They would take it in a heartbeat.
Ben Lorica. You suspect they would use that?
David Fattal. Of course. They would take it in a heartbeat. Stereo is the same content as video, but with the truth about the geometry of the scene. You don’t have to work harder.
Maybe another way of saying this is that I always say one stereo image is worth maybe 1,000 regular images, in the sense that perhaps you could compensate with volume, and that’s what people are trying to do. But if that data were available, you wouldn’t have to work so hard to extract reality from it. You would get it for free, even in real time, and that can help boost the efficiency of these models even further.
There are a lot of advantages. As you mentioned, a lot of people and companies are going through the trouble of wearing VR headsets and training robots. By the way, they have binocular vision in that case because they need to thread the needle. A lot of these VLAs, for example, which correlate spatial understanding and actions, are trained by operators wearing headsets today and threading the needle. That’s stereo data. It’s just very forced.
With video, you have the whole corpus of YouTube and Instagram. Essentially, the entire web is already there, so you don’t have to do anything. It’s very expensive to pay people to acquire data, and it’s always going to be extremely niche. It will be robots doing one type of action, or maybe you’re going to hire somebody to go to an airport or a warehouse and get robots to operate there.
But I think what we’re talking about is training general-purpose intelligence. For that, you want to facilitate the capture of a corpus that’s analogous to what we have on Instagram and YouTube, which is crowdsourced. It’s just people randomly taking 3D pictures of their environments. It’s very diverse, everywhere in the world, in every condition, with all the messiness of the real world.
There’s smoke. There’s a ray of light that’s going to hit one camera and not the other. You might be underwater. You might be in a very weird position. You might have a sensor that’s defective. You might be moving and have motion blur. All of that is extremely relevant, and you want to have that type of data to augment the training set for our models.
Ben Lorica. And obviously, there are things that are hard to model, period, like turbulence.
David Fattal. Yeah.
Ben Lorica. A couple of questions. First, let’s say I’m sold on stereo data, but I happen to have mostly monocular video. Are there people trying to build models that recover some of this information just from monocular video? In other words, I have monocular video and I’m going to build a model to enrich it so that it feels like stereo. Is that possible?
David Fattal. Yes. First of all, some people are trying to do everything just with monocular video. For example, there’s a company called Roda AI, I think, in the Valley.
Ben Lorica. Yeah, I talked to them.
David Fattal. You talked to them, right?
Ben Lorica. So Roda AI isn’t even going to bother to say, “Okay, since we have a lot of video, we’re just using video.” They’re not even going to try to recover information that might exist that could turn this mono video into stereo?
David Fattal. That’s my understanding. I think we’d need to confirm with them. Maybe they have some research, but based on whatever is public, I think that’s their whole credo.
Ben Lorica. Let’s say I have mono video, but it happens to be the case that there were a few cameras involved in filming the same scene.
David Fattal. Yeah. The first obvious case is that if I move my camera over time, you’re essentially getting several points of view of a scene. This is why 2D videos already have the ability to give you some kind of geometry. It’s called structure from motion.
Let’s say the scene was perfectly still. If I’m going to move around with my camera, it’s like having a multiview image, and then you can reconstruct the geometry. That’s basically the technique.
So a 2D video does contain a lot of information about the structure of the world. The problem is you’re mixing what we call parallax, which is frozen time and moving the camera, with the time axis.
If you try to do structure from motion on a waterfall, for example, it doesn’t work. As you move, the water falls. What you’re trying to dissociate is the time axis from the parallax, or essentially the 3D view axis. That’s the level of complexity that comes on top.
This is where the messy environment comes into play. If I have a smoky environment, a cloud, or something else that moves at the same time I’m trying to figure out the geometry, moving the camera might not be enough.
I think the consensus among experts is what I said in the beginning. You can use 2D videos for most of your training. You can also synthesize 3D or stereo. You can use a rendering engine like Unity or Unreal and create a scene, then render all the points of view you want. It’s all perfect.
But again, you don’t capture what I call the long tail. The long tail is all these messy situations that you can’t think about. For that, I think you don’t need nearly as much real stereo data, but even if you had a million of these stereo pictures for a billion regular mono pictures, I think that would go a long way toward fixing the last mile and getting the models to be perfect.
Ben Lorica. One of the favorite hacks of anyone who’s starved for data these days, including the people building robotics foundation models, is synthetic data. What’s the state of synthetic data generation for stereo?
David Fattal. Again, it depends on who you talk to, and a lot of the time it’s motivated by people’s own needs, whether to fundraise or otherwise.
NVIDIA has a wonderful Cosmos model that’s basically a really good 3D renderer. They can model all kinds of complex lighting effects, and they have products where you can start to simulate your robot and generate purely synthetic data.
If you listen to NVIDIA, that might be enough. It might be the case that for a very specific scenario, with a certain robot training in a known environment, that might be enough. It’s not going to start raining or get foggy in your warehouse.
But again, I go back to the goal of training some kind of general intelligence that understands space like humans do. If you want to get to that point, my contention, and that of a lot of people in the field, is that you need that long tail of stereo.
The stereo or multiview imagery generated from these rendering engines is beautiful, but you can ask anyone to do a test. Take a real picture of the real world, and then take the most beautiful render from Unreal Engine on a super-expensive, latest NVIDIA GPU. Humans, at least so far, can always tell that one is synthetic and the other is real.
There’s something about the real world that’s not so perfect and that you can’t necessarily pinpoint, but you’re definitely able to tell. It’s that element that you want to bring into the training.
Ben Lorica. By the way, even with plain foundation models, many companies are coming to realize that data is actually the main asset and moat.
You can rent some things. You can think of the model as generic and swappable, but what will distinguish you as a company is your ability to leverage your own data, either to specialize the model or to capture that compounding value where, as people use your model, you capture more data, and therefore your entire workflow gets better.
That’s just for the simpler foundation model. I can imagine that in the scenarios we’re talking about, it’s even more pronounced.
David Fattal. Yeah, exactly. There’s a lot more capacity to ingest and use all of the small details I mentioned. They’re not obvious to us, but there’s obviously a lot more information in pictures than in text. It’s much more complex.
But I think you’re right. The only metric for these models, even the simpler ones before, was the number of weights. You were just counting the number of weights. You went from millions to 10 million, 100 million, billions, and so on.
I think people are starting to quote the amount of data that’s going in, and that’s a different dimension. And particularly, their own data.
Ben Lorica. I think people are realizing that, to some extent, with foundation models we’re getting to the point where, okay, some are better than others, but the differentiation is beginning to become less and less. Really, the differentiation is your data.
David Fattal. Yeah, absolutely. I mean, we’ve run out of public data, so that’s basically what it is. The only difference that you can make now is your data.
And by the way, you can use distillation, as many people do. You can take a very expensive model and, with maybe $1,000 or $10,000 worth of training, distill 90% or 95% of it. Unless you have a data moat, I think all of these models are going to get pretty much the same.
Ben Lorica. In the case of Immersity, it seems like the bet is that you have these devices, so you almost have a device-data flywheel that can potentially happen.
How exactly does it work? People can use their own phones to start capturing this rich data. Is it a separate app on the phone? In other words, is the hardware on the phone already there to capture this kind of stereo data?
David Fattal. Yeah, correct. What we fundamentally do at Leia is develop technology to turn any screen into a spatial screen, or a 3D screen, meaning the screen is able to project two different stereo views and send them to your two eyes, which a regular display doesn’t do.
Ben Lorica. How different is that user experience, David, from wearing one of these expensive Apple headsets?
David Fattal. Wearing a headset is essentially like strapping two displays in front of your eyes. It’s completely immersive. You have 360-degree, complete immersion.
With our technology, you get most of it. Imagine that your laptop transforms into a box in front of you where content is rendered. Of course, it’s not completely immersive. If I don’t look at my screen, I’m still in the real world, which has a lot of advantages. I can still talk to my friends. I can still interact.
But when I look at the screen, it presents information and content as a volume.
Ben Lorica. Obviously, the people who are generating content for the web are generating content for plain screens. Does that have an impact? Apple had to purposely generate content for their hardware, right?
David Fattal. Maybe 10 years ago, it was true that it was very difficult to create 3D content. I think in 2026 people don’t realize that 3D content is everywhere, even though you think it’s 2D because the only way to experience it is on a 2D screen.
The 3D content is plentiful. You just mash it onto a flat surface before we consume it. The whole point of the technology, and what we’ve been doing at Leia for the last 10 years, is saying that 3D content is meant to be experienced in 3D. Therefore, let’s create technology that’s going to allow any device, a phone, tablet, laptop, or computer, to express that content the way it was meant to be experienced.
Take any game today. Any game that has been created with Unity or Unreal Engine, literally any game for any console, is 3D behind the scenes. It’s essentially 3D objects, 3D meshes, textures, and so on. The reason you see it in 2D is because the game has a virtual camera filming the scene and showing it to you, and it just has one camera.
If you have the ability to show different content to your two eyes, all you need to do is get in there and add a second camera to the scene. You don’t need to re-author the content or redo the expensive part of creating the characters, shading, lights, and all of that. That’s already done. All you’re doing is saying, “Hey, I’m coming in with a second eye, and I’m re-rendering the scene from two points of view.”
The other thing that’s really little known, and I think you alluded to it, is that anybody owning an iPhone, a Samsung phone, or one of the flagship phones these days has these two cameras. If you go to your camera app and scroll through a couple of options, you’ll find a spatial recording option.
What does the spatial option do? Of course, it’s Apple, so they call it a different name, but it’s just good old stereo. It takes two pictures, puts them side by side, and you get a stereo picture.
In that sense, billions of people today have a device in their pocket that could be recording the world in 3D. They could record everything they do, their kids, their pets, their travels, the same way that they do today and then post to YouTube or Instagram.
The reason you don’t do that, and I don’t know if you have an iPhone, but even I don’t use the spatial feature, is because there’s nowhere today to consume or visualize it. The only way you can visualize your spatial content is to buy that expensive piece of equipment. You need to buy the Vision Pro. It’s at home, it has the battery pack, and you need to set aside time to put it on. It’s just too complex.
Now imagine that the iPhone had a 3D screen. Every time you took a picture, you’d get the instant gratification of seeing the scene in an immersive way. That would completely change the equation.
As a matter of fact, at Leia we had a project a couple of years ago. We were working with a company called RED Digital Cinema, and we had this RED Hydrogen phone that came out around 2018. It had a stereo camera in the back and in the front, and it had a 3D screen.
It was still early in the technology’s development, so the 3D screen wasn’t nearly as good as it is today. But I think it was already very compelling. The number one thing people were doing on that phone wasn’t watching movies or playing games. They were taking pictures of the world and posting them.
We had a social media app called LeiaPix. They were sharing and creating this community of people who were rediscovering the world through these 3D pictures. The reason it took off was that every time you took a picture or video, you were able to see it in 3D right away.
That’s what we’re trying to promote at the company. These types of screens are coming. The reason they’re coming is because, after 10 years of development and research, they’re getting to the point where they’re extremely good. They’re very comfortable to look at, and the price point is coming down as well.
In a couple of years, it will cost about the same to upgrade any screen with 3D as it does to upgrade a screen with touch. Today, on your phone, you don’t even think twice. Of course there’s touch. Every phone has touch.
For us, and for a lot of people in the industry, the same thing is going to happen with 3D technology. In a couple of years, it’s going to be unthinkable not to have this 3D add-on, because it doesn’t take anything away from your regular experience. Just like touch doesn’t degrade your image quality, adding 3D with our technology doesn’t degrade the 2D image quality.
You can still experience all your apps in Retina resolution and so on. You just have the option to turn on 3D. When you want to take a 3D picture, you take it and then you can experience it.
If you believe in that future, what should happen is that people are naturally going to start taking these stereo pictures. They’re going to start posting them. Instagram and YouTube are going to support them. They used to support these 3D formats, and they don’t anymore, but they can again.
That’s going to create that flywheel. More people take more stereo data. You can improve your world models. Then you can start creating better games, better experiences, better AI. That, in turn, gives you better devices, and that flywheel gets going.
Ben Lorica. Does that mean, David, that if you fast-forward to the future you just described, most of the user-generated content on, say, YouTube will naturally be 3D?
David Fattal. Yes.
Ben Lorica. Then that means the foundation model builders who are building foundation models for robotics, or whatever the use case is, benefit from that.
David Fattal. Exactly. Everybody will benefit.
In our view, it’s a faster way to get to ubiquitous stereo data than relying on, say, self-driving cars. Self-driving cars can take multiview data, but it’s contained to the road.
I think Mark Zuckerberg’s vision is also that everybody is going to wear these AR glasses, which, by the way, have stereo. That’s another way to get there. If somehow there are a billion people in five years wearing glasses and recording the world, that’s another way.
Our contention is that it’s much more likely you’re going to upgrade existing devices and give people a reason to capture the world in stereo, because they get the immediate gratification of seeing the content on their device. That will get us to the data collection a lot faster than any of these other methods.
Ben Lorica. It sounds like what you’re saying is there’s no software upgrade. It has to be a hardware upgrade. In other words, our laptops and phones have to have a certain type of hardware to consume this content.
David Fattal. Yeah. I think you can fake immersion to some degree by having parallax. If I track my head, I can move the content.
I don’t know if you’ve ever experienced this on Facebook. They have this type of 3D effect. Actually, even the iPhone has it now. They give you a preview of your picture, and when you tilt it, it gives you a little bit of motion.
But once you experience one of our displays, having true, different images in your two eyes is what creates the sense of volume. Even when you don’t move your head, you basically have an impression of feeling the space that you don’t get with these other methods.
These other methods are useful because people are getting used to converting a still picture into something that looks like it has depth. Once you have a lot of this content and then have the proper 3D screen, you can really express it in a volumetric, immersive way.
But in order to really benefit, to get the satisfaction and sense of presence, you need these two eyes. So some kind of hardware upgrade is required.
Ben Lorica. What’s the timeline for a company like Apple to push this onto its phones, tablets, and laptops?
David Fattal. Again, I don’t want to talk for any particular company. Apple is perhaps the most difficult case because they tend to be second movers. They’re never first. They like to observe what’s going on in the industry.
Ben Lorica. I guess the Chinese will go first, right?
David Fattal. I think your guess is right. Some Chinese companies will go first. We talk to a lot of customers, so of course we see a little bit further into the future. It’s very likely that a Chinese company will come first.
Ben Lorica. I guess two questions. One, give us a timeline for when a Xiaomi, Oppo, or whoever might push it out. Secondly, what would be the additional cost of the device for the consumer?
David Fattal. What we see today is that the most interested parties are the monitor and laptop companies, because there are some immediate applications.
There are successful gaming programs. We just had a project with Samsung, for example, the Odyssey 3D, that was purely focused on gaming.
Ben Lorica. So the idea is that there will be certain laptops or desktops targeting certain slices of the gaming market?
David Fattal. Exactly. Certain verticals.
Ben Lorica. What would be the guesstimate timeline, then?
David Fattal. It’s already out. Samsung came out last year with the Odyssey 3D for gaming, and we had a really successful program.
Ben Lorica. What was the uptake?
David Fattal. As far as I know, they sold everything that they ordered from us.
Ben Lorica. How much more expensive was it? Ten percent?
David Fattal. I think it was retailing at $2,000 when it came out. Usually, you compare this to a high-end OLED. You don’t compare it to the low end. I think the OLED was probably about $1,800 and the 3D was $2,000, so it’s about 10%.
Ben Lorica. Slightly more than 10%.
David Fattal. That’s right.
Other verticals are also proving successful. One is with Barco. They partnered with a company called Avatar Medical, which has the first software that’s been FDA-approved to be used in the operating room. That was a big win. In medical applications, obviously, 3D is very interesting.
Then you have telepresence. You have Google Beam in partnership with HP, another device that’s really high-end. For telepresence, it’s a beautiful experience. I think if you get the chance to experience it, you’ll see.
Interestingly, it’s moving from bigger devices toward smaller devices. It’s anybody’s guess when one of these companies will really want to move seriously. Again, I can’t reveal anything that’s private.
But I think it’s fair to say that there’s been a pause because of AI. If you’re an OEM, if you’re Xiaomi, Oppo, or Samsung, today your entire energy is devoted to running AI, either as a cloud service on your phone or on the edge, so you can erase the background and do all of these things.
They only have an appetite for one thing at a time. For the past two years, it’s been AI. But 3D has always been on the radar. My guess is that within three years, you’ll probably see the first real volume product from one of these vendors hit the market.
Ben Lorica. So the people building these foundation world models, or foundation models for robotics, are probably chomping at the bit. They want this data, but the rollout of the hardware is slow, so they have to build their models on the current data they have.
David Fattal. Correct. But some of these companies, and again, the Chinese companies are pretty aware of this, have the funds and bandwidth to think a little further ahead and say, “Okay, maybe we’re an AI company or a content company, but having that hardware program might really accelerate data collection, and therefore it’s very useful.”
You’ve heard the rumors about OpenAI having some kind of hardware device with Jony Ive and so on. It might not be for 3D, but it’s for the same purpose. They’re not doing this to be a consumer hardware company. They’re doing this because it’s a collection device that will help with data collection.
If you understand the technology and you want to think a little bit ahead, and you’re one of these big tech companies, you would definitely consider that thesis. You might need, or will almost certainly need, new hardware to reinforce AI training.
Ben Lorica. Is it fair to say that 3D TV failed in the past?
David Fattal. Oh yeah, absolutely.
Ben Lorica. Then what’s the guarantee that this time is better?
David Fattal. That’s a great question. We created the company in 2014, exactly when 3D TV started to fail, so I experienced it firsthand. We had to raise money in that environment.
I think 3D TV was way ahead of its time. There was no content. There was, by the way, no AI to generate content or help you create 3D content. Everything had to be painfully filmed with a stereo camera. It was super expensive.
Of course, if you’re James Cameron and you can film Avatar in 3D, that’s different. That’s pretty much what launched the 3D TV movement. Avatar came out in 2009, and people thought there were going to be a lot more movies like it.
And of course, you had to wear glasses as well. Having to wear an extra layer of anything on your face is just not natural for people.
If you look at today, the comfort is undeniable. You can simply look at the screen, and the 3D can disappear. You’re not buying a special device that forces you to consume technology in a particular way, like having a headset, or something that decreases your resolution.
Your device is going to look the same. You just have the option to turn on 3D. I think that switchability aspect is the fundamental difference between having a single-trick pony that you use maybe once and then put in a drawer, and having something that’s truly useful and always with you.
That’s the feedback we’ve been getting from the OEMs, and that’s always where we start. For us, 2D is sacred. You cannot compromise on 2D quality.
Then we try to add 3D in the least invasive way. It doesn’t require more power. It doesn’t degrade your 2D experience. It’s almost the same cost, and so on. We’ve worked very, very hard to make it happen this way.
Ben Lorica. As you alluded to, there are definitely domains and sectors where this will resonate quickly. Medical is one that you mentioned. I imagine the military is another sector. There are a bunch of industrial sectors where this might work.
But for mass-scale data gathering, you almost want some consumer company to take the lead. The ones that come to mind are YouTube and TikTok.
In the YouTube case, you can imagine Google, because they own YouTube, coming out with a device ahead of time. TikTok might be more of a follower because once the devices are out there, they can turn on support and make that available. Otherwise, they would have to create their own devices, right?
David Fattal. Yeah, but I wouldn’t count TikTok out. TikTok is ByteDance, and I remind you that ByteDance acquired Pico, which is a headset company.
So the two players you mentioned are right on the money. I think those are two of them.
Ben Lorica. What’s the future of the dedicated headset in this world that you’ve laid out?
David Fattal. I think the headset is extremely good for completely immersive gaming, for sure. It’s very different to have a phone or a laptop versus being completely immersed.
First of all, your hands are free. You can hold your weapon. You can start to grab things instead of having to type on your keyboard and so on. So for certain types of completely immersive experiences, or if you want to be trained in a warehouse environment, it makes sense.
Ben Lorica. But Apple was touting it for watching movies, and I think that’s overkill, right?
David Fattal. I think that’s overkill. Watching a movie is the exact example where the 3D display on your device is perfectly suited.
Where do you watch movies today? If you look at younger generations, they’re all on their phones. On the bus, on the train, they’re watching on their phones. They don’t even watch TV anymore.
You want to go where the consumer is today. You don’t want to have to reeducate the consumer. You want to adapt yourself to where the consumer already is. That’s what we’re trying to do.
Ben Lorica. We’ll put you on the spot and make you predict. One year, two years, three years. What milestones are you looking for to signify that what you’ve laid out is starting to happen?
David Fattal. I think within one year you’ll see one more major OEM, besides Samsung, take on the technology and put it on either a monitor or a laptop, with some significant volume.
I think within two years you’ll see some of the content companies, maybe Zoom or Teams, support telepresence apps in stereo, or some of the gaming companies start to support rendering their games in stereo very naturally.
And I think within three years you’ll probably have the first high-end mobile device with a 3D screen that can do all of the capture and visualization in 3D.
Ben Lorica. By the way, I’m going to take advantage of the fact that you’re an expert on spatial intelligence. Where do you sit, David, in the self-driving car debate: camera only versus lidar?
I’m much more in the sensor-fusion, multiple-sensor, multiple-levels-of-redundancy camp. I want to be able to see in fog and rain, rather than the Elon camp of, “I’m just going to use cameras.”
David Fattal. Unfortunately, this has turned into a political issue of whether you like Elon or not.
Ben Lorica. No, no. I’m just thinking in terms of common sense, because I want my self-driving car to see under any conditions.
David Fattal. Yeah. The argument for common sense is actually Elon’s and Andrew Karpathy’s argument, which is that we’re human and we drive with two eyes. Whether it’s fog or something else, we’re able to drive the car. Maybe we have accidents, but we’re human.
Ben Lorica. Yeah, but why not take advantage of super vision?
David Fattal. You should do it at least as well as humans. I think you’ll be able to do everything that you need with cameras. I’m more in Elon’s camp. I think you’ll be able to do everything that you want to do.
Ben Lorica. What about heavy rain?
David Fattal. Sure. If you want to see through conditions like that, or see through walls or things like this, then other sensors can help. I think it’s just more expensive.
Ben Lorica. But lidar is coming down in price, right?
David Fattal. Yeah, it’s coming down in price. But it looks a little bit silly on top of your car. Look at a Tesla and look at a Waymo, for example.
Ben Lorica. Look at the Cybertruck. That’s not exactly beautiful either, right?
David Fattal. That’s not because of the sensors. That’s a pure design decision.
Ben Lorica. All right. Closing. This is a bit of a left-field question, but for our listeners who want to know more about these types of technologies and how they may play out in the future, are there science fiction books that you recommend that capture this well, or books in general that you recommend?
David Fattal. I think the classic one is Minority Report. It’s not a book. It’s the movie with Tom Cruise where you have this transparent world and he interacts with it. It’s actually maybe the closest to what I want to leave you with in closing.
Think of the content. Content is king. The 3D content and the spatial relationships between objects are king, and then you have different ways to experience it. You might want to put on a headset, or it might be through a mobile terminal or a laptop.
Eventually, what you want is to interconnect these things. You want to have this immersive space where people can tune in and out, either from a headset or a phone, at any moment. They can interact, do some work, have some fun activities in there, and then tune out.
It includes the 3D rendering, but it also includes the interaction and so on.
If you rewatch the Minority Report movie now, having listened to the podcast, I think it’s going to make sense. It’s a little old. Tom Cruise needs to wear special gloves, for example, to interact. You don’t need that anymore. Computer vision can very well tell all of your fingers apart, so you wouldn’t even need that type of equipment.
That would be my pick if you want to get a taste of what’s to come.
Ben Lorica. And with that, I guess this conversation is a classic example of that old saying: “The future is already here. It’s just not evenly distributed.”

