Why video is the first frontier as AI learns to read the physical world


Photo of AI Learning and Artificial Intelligence Concept. Business, modern technology, internet and networking concept.

AI Learning and Artificial Intelligence Concept.

Image Credits Credit: Canva

TL;DR

Hundreds of millions of cameras are already deployed, but most footage still requires a human to know what to look for. Lumana, founded by ex-Intel computer vision leaders, processes over a billion images daily across 50,000+ cameras using its VIA-1 model, which learns what is normal for each individual camera and flags deviations. The company filters locally before sending anything to the cloud, following a principle of “filter before you spend.” Video may be physical AI’s natural starting point because the infrastructure already exists.

A camera overlooking a loading dock might record twelve hours of trucks arriving, workers moving through the site, and boxes leaving the building. Most days, nobody has a reason to watch any of it. The footage only becomes useful when a package goes missing, an accident happens or somebody needs to work backwards from an event and find out what happened.

That has been one of the strange limitations of video surveillance for years. Cameras became digital long ago, but the footage they produce still depends heavily on a person knowing what to look for and where to find it. AI is beginning to make more of that footage understandable and searchable while events are still unfolding.

Axis Communications estimates that 562 million surveillance cameras were installed worldwide outside China by the end of 2025. More of those cameras are also arriving with intelligence built in. About two-thirds of cameras shipped in 2024 included deep-learning analytics, according to the company’s research.

For companies building what is increasingly called physical AI, much of the infrastructure is therefore already hanging from walls and ceilings. The opportunity is in making those existing cameras more useful by teaching software to understand what is happening in front of them.

Lumana is one of the companies betting on that transition. Its founders came to the problem with years of experience in computer vision at Intel. CEO Sagi Ben Moshe previously led Intel’s RealSense business, while CTO Ofir Mulla worked on the architecture behind its 3D and LiDAR cameras.

The California-based startup raised $40 million in July 2025 in a Series A led by Wing Venture Capital, with participation from Norwest Venture Partners and S Capital, taking its total funding to $64 million. By December, it reported that more than 50,000 cameras connected to its platform, with customers including Fortune 500 businesses across the United States.

The company says its AI video surveillance systems now process more than a billion images a day across more than 50,000 cameras. For customers, however, the early changes are usually much more ordinary than that number suggests.

First, stop staring at the screens

For decades, a familiar sight in security control rooms has been a wall of video feeds and people expected to notice when something looks wrong.

Photo of Ofir Mulla, CTO of Lumana
Ofir Mulla, CTO of Lumana — Credit: Lumana

Lumana CTO Ofir Mulla says teams spend the first month after deploying its technology mostly sorting out much more ordinary problems. Teams find cameras that are offline, work out where coverage is inconsistent, and decide who should receive which alerts.

As that work settles, operators can spend less time trying to follow every feed and more time looking at the events the system has flagged for investigation.

Instead of spending time watching passive video walls and trying to spot something unusual, operators can focus their attention on events that warrant investigation,” Mulla said in an interview.

Doing that well depends heavily on context, something traditional motion detection and fixed rules have struggled to capture.

A person standing beside a warehouse door might be perfectly normal at 2pm and worth investigating at 2am. A delivery truck parked at a loading dock for 20 minutes could be expected. The same truck sitting there for three hours may not be.

Lumana’s VIA-1 model is designed to learn from the environment seen by each camera rather than apply exactly the same definition of normal everywhere. The company says this can reduce false alerts by up to 90% compared with legacy motion detection and rule-based systems, although Mulla is careful to describe that figure as the higher end of what Lumana has observed rather than a result every customer should expect.

But the much harder test comes when the environment itself changes.

What happens when normal changes?

What looks normal to a camera can change quickly. A warehouse might rearrange its layout over a weekend, while a retailer suddenly has to deal with the Christmas rush or a factory introduces a night shift that brings dozens of people into an area that used to be empty at that hour. In each case, activity that might have raised an alert yesterday could be perfectly routine today.

Mulla says VIA-1 can adjust its understanding of an individual camera as its environment changes, with feedback from operators helping when something has materially shifted. He would not give a fixed time for how long that adjustment takes, saying it depends on the scale and type of change.

A useful AI video system has to learn that a new shift pattern is now routine while still noticing the activity that should raise questions. Because the time needed to adjust depends on what has changed, operator feedback remains part of that process, particularly when a site has undergone a substantial change. In practice, the system is learning alongside an environment that does not stay still for long.

Making all of this footage easier to search and understand also raises privacy questions, particularly in workplaces and other spaces where people are routinely recorded. Lumana allows customers to set retention and access policies and disable features such as face and gender recognition depending on local requirements. But as existing cameras become more capable, companies also have to consider what they should collect, who should be able to search it, and how long that information should be kept.

Today, much of the work is still about helping people decide where to look. Lumana sees a larger role for the same technology as these systems become more capable.

Mulla describes “physical AI agents” that can divide up monitoring, verification, and response. Lumana’s platform already supports actions ranging from sending notifications and triggering webhooks to activating connected systems when particular events are detected.

That puts the camera in a different place inside the business. Footage can become an input to software that interprets an event and, in some cases, triggers what happens next.

Filter before you spend

Doing this across tens of thousands of cameras gets expensive quickly because video is costly to move and process at scale. Sending every frame to the cloud for more intensive AI processing would rapidly become expensive, particularly as camera resolution and deployment sizes grow.

Lumana handles much of the continuous processing locally instead. Its Core hardware sits close to the cameras and performs the initial processing there. According to the company’s platform overview, most video processing runs locally, while the cloud provides additional processing, remote access, and management across sites.

The vast majority of video is routine and never needs to leave the site,” Mulla said.

Only footage that warrants deeper analysis needs more expensive processing. The company’s principle is more simply: “We filter before we spend.

The same economics are likely to matter well beyond Lumana. Cameras and other sensors generate too much information for every frame to receive the same amount of computation, which makes deciding what deserves deeper analysis part of the technical and economic problem.

Video also has one advantage over some of the more ambitious ideas around physical AI. Companies do not have to wait for millions of new robots or rebuild their sites around new hardware. Hundreds of millions of cameras are already pointed at factories, shops, universities, warehouses, and streets.

Most of those cameras were installed to record what happened. If AI can reliably understand more of that footage as events unfold, the same infrastructure can help companies understand what is happening while it is still happening and, in some cases, decide what should happen next. And that might just be what makes video such a natural starting point for AI’s move into the physical world.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top