Seeing the Future
How the Team Behind Kinect and Face ID Is Evolving Perception for Physical AI
January 5, 2026
In 2005, Alexander Shpunt co-founded PrimeSense to work on a problem that would define the next two decades of his life: how do you teach a machine to see depth?
Not capture images. See. Understand that the wall is ten feet away and the person is three feet away and the hand reaching toward you is closing the distance. Perceive space the way a human does. Not as flat pixels, but as dimension, distance, the relationship between objects in three-dimensional space.
Alex's solution used light. An infrared projector casts invisible dots across a scene; a camera reads how the pattern distorts at different distances and triangulates it into a real-time depth map. They called it Light Coding. It was inexpensive, fast, and could be miniaturized onto a chip.
Five years later, that technology powered Kinect. Microsoft sold eight million units in sixty days. Guinness certified it as the fastest-selling consumer electronics device in history, outpacing iPhone and iPad in their launch windows.
Living rooms all around the world now had a simple, but powerful, sensor bar above their TV, ready to play. The machine sees you, and moves with you, not as a 2-D silhouette, but as a 3-D skeleton. It tracks your arms, legs, torso, head. It knows where you are in the room and how you're moving. No controller - your body is the interface.
Every Kinect saw the same way with millions of households and millions of players perceived through identical eyes. When you played online with someone across the world, both of you were being seen by the same depth-sensing architecture. Shared perception at scale. That's why developers could build for it. That's why it became a platform.
Apple acquired the company in 2013, and the PrimeSense team kept building. The same core discipline evolved to map faces across small screens, within arm's length from the user.
Face ID shipped in 2017 and is now in billions of Apple devices. Thirty thousand infrared dots projected onto your face, read by a camera, triangulated into a 3D depth map that confirms you are you. As simple as that - your face in full depth.
And again: every iPhone sees faces the same way. Billions of devices, shared perception. That's why it's trusted, and that's why it works at scale.
MIT named the underlying technology one of the most important of its era. The World Economic Forum called the team Technology Pioneers. The depth-sensing architecture they'd invented now lives in pockets and homes across the planet, touching billions of lives daily.
By 2021, artificial intelligence was crossing a threshold. The models that had learned to read text and recognize images were no longer confined to screens. They were being asked to navigate warehouses, operate machinery, share sidewalks, and roads. AI was becoming physical and present in the world.
The team that shipped perception technology in billions of devices recognized the pattern immediately. They had seen this moment before. But this time, the stakes were different. A smartphone that occasionally misreads depth is an annoyance. A warehouse robot that can't see a worker in its path is a catastrophe.
And across the industry, a troubling pattern was becoming visible. Perception was often approached as a parts problem rather than a systems one. Many robotics teams assembled perception from components—different sensors, different vendors, different calibration stacks. Each machine ended up seeing the world through the specifics of its own build, its vision stitched together rather than coherently designed. The result was fragmented world models. And when machines do not share a consistent understanding of the world, reliable and safe operation becomes fundamentally harder to achieve.
So they started building again.
The Fourth Dimension
Structured light had worked for rooms and faces. An infrared pattern projected, read, triangulated. But structured light has limits. It works at close range and only captures where things are, not where they're going.
For machines that move through the world, that's not enough. A robot navigating a warehouse doesn't just need to know where the forklift is, it needs to know the forklift is moving toward it at four meters per second. A delivery bot on a sidewalk doesn't just need to see the child, it needs to see the child is running.
Traditional sensors capture position. To understand motion, software compares frames. Position now versus position a moment ago. This introduces lag. In a world that moves, lag is where failures happen.
The team had solved this before. Kinect didn't just see where your body was, it tracked how you were moving. Face ID doesn't just map your face, it confirms the face is live, present, real. Both technologies understood scenes as they evolved, not as frozen moments.
Now they needed to do the same thing at range, at speed and for machines operating in open space.
The answer was coherent vision. A different way of using light. Instead of projecting a pattern and reading distortion, coherent vision sends a continuous signal and measures what returns. Position and velocity captured in the same instant. Not calculated afterward. Native to the measurement.
The fourth dimension is velocity. Where something is, and where it's going, known together.
A Prototype Is Not a Solution
But having the technology isn't enough. This team had learned that lesson shipping billions of devices. They knew what survives the transition from lab to field, what a design decision costs when it's multiplied a million times, and the difference between a prototype and a product.
A demo that works on stage is not a solution that works at scale. The industry was full of demos. So they did what they've always done: designed it to scale from the start.
Custom silicon, with photonics and electronics integrated at the chip level. A coherent 4D sensor fused with RGB and motion awareness in a single architecture. Environmental robustness and interference immunity designed together, not bolted on afterward.
Fabrication by the same foundries that supply the automotive and consumer industries. Advanced packaging executed by the world's leading semiconductor assembly provider. Design choices, interfaces, and calibration strategies defined from the start with manufacturing in mind.
The result: a single module with one connector. Multiple sensors, unified output. Position and velocity captured in real time, ready to deploy. One module, one way of seeing—so every machine that uses it shares a consistent view of the world.
They didn't set out to solve integration. They set out to build perception that scales. But at this level, the two become inseparable.
Integral Simplicity
The module is only the beginning. Perception has to go somewhere—eyes connect to brains through the nervous system, and a robot that sees the world still has to understand it. That means connecting sensors to silicon to software to AI compute: data flowing from the edge to the cloud and back, models processing what the machine perceives, making decisions, issuing commands. The full loop, closing in milliseconds.
Today, this is where robotics companies drown. They solve the sensor problem and discover another problem behind it: How do you format the data? How do you manage latency? How do you connect to different compute architectures? How do you pipe perception into the AI stack without building a custom integration for every deployment?
It's not one integration, it's integrations on top of integrations. Sensors to processors, processors to networks, networks to data centers, data centers to models. Every connection is a potential failure point. Every handoff hides latency and incoherence.
Lyte doesn't just unify the sensors. It unifies the path from perception to intelligence: sensors fused to silicon designed to talk to software, software designed to talk to AI compute. The full stack, from photons hitting the sensor to decisions coming back to the machine - one architecture, no seams.
The complexity doesn't disappear, it gets absorbed. What remains is simple: a perception layer that connects to the intelligence layer. Eyes that talk to brains.
That's what makes shared understanding possible—not just machines that see the same world, but machines that can think about it together.
Perception, Evolved
Robots are entering homes and working in warehouses and hospitals to deliver medicine and move inventory. They're sharing our roads and sidewalks, transporting people and products. We're entering an era where machines won't just compute. They're learning to perceive, decide, and act. Not just capture data, but understand space and read motion. They're operating alongside people in a world that doesn't hold still. The margin for error is gone.
For twenty years, this team has been solving one problem: how do machines see deeply? From living rooms to faces to the open world. From millions of devices to billions. Each chapter larger than the last.
Now the problem has reached its full scale. Machines that share our world must share their understanding of it. They must see together. Know together. Act together.
The team that gave sight to game consoles, then phones, now gives sight to physical AI.
Introducing Lyte.
Perception, Evolved