← Research Notes
Oct 09, 2026·Simam Digital Research·Reviewed Oct 09, 2026·7 min read

What is spatial intelligence in AI? A working definition from the lab

Spatial intelligence is AI that understands where things are, how they relate in three dimensions and how they change over time, and can answer questions or act on that. Here is what it can do today, with numbers.

Spatial intelligence in AI means a system that understands space, not just pictures. It knows where things are in three dimensions, how big they are, what is next to what, how a place changes over time, and it can answer questions or take actions based on that. An image model can tell you there is a tree in a photo. A spatially intelligent system can tell you where the trees are relative to the tower, which side of it faces the park, and that there are about 3,900 square metres of trees on the site.

The phrase has two lives. In psychology it is one of Howard Gardner's multiple intelligences: the human ability to picture and reason about space. In AI it has been popularised by Fei-Fei Li and World Labs as the next step after language models, the subject of our note on Atlas and world models. This note is about the AI meaning, from the point of view of a lab that builds with it every week.

Three abilities

We find it useful to split spatial intelligence into three abilities.

  • Reconstruct: turn images or video into a 3D model of a real place. In our apartment viewer, a 29-second drone clip became a 3D site in 9 minutes 37 seconds on one desktop graphics card.
  • Understand: name and find things inside that 3D. The same site can be searched for trees, green space, windows, roads, the rail line and cars, with every part of the model shaded by how confident the AI is (see how that search works).
  • Measure, reason and act: answer questions in real units, plan, and simulate. From one ordinary video, using the people in it as rulers, we measured a forklift passing 2.45 m from a walking worker. From a 10-second phone walk processed in the browser, the camera's path came out within 2.7% of a full photogrammetry solve.

How it differs from computer vision

Classic computer vision labels pixels in a single frame. Spatial intelligence keeps a consistent world across many frames and viewpoints, so an answer does not change when the camera moves. That consistency is what makes measurement possible, and it is also what lets you put information on the world, such as a building's height, a plot's owner or a unit's availability, and have it stay on the right floor as you orbit.

What it cannot do yet

  • See what the camera never saw. A drone that flies one side of a building leaves the other side soft or missing. A fixed camera on a street holds together only near its own line of sight.
  • Know scale on its own. One camera gives shape but not size. We set scale from counted storeys (±12% in the apartment viewer) or people's height (11 to 13% agreement on London streets). Survey-grade work still needs surveyed points.
  • Tell captured from imagined. World models can fill in unseen parts convincingly. That is useful for design and storytelling, but a measured twin has to label what was observed and what was generated.

How to judge a spatial AI claim

Ask for numbers on held-out views (frames the model never trained on), real units with stated error, the time and hardware used, and a list of what went wrong. We publish all four for every experiment on the lab bench, and we explain the reasoning in how we evaluate spatial AI workflows. That is the practical test of spatial intelligence: not whether it looks impressive, but whether its answers hold up in the place itself.

Source references
Business relevance

Most operational questions are spatial: where is it, how big is it, what is next to it, what changed. AI that can answer those from ordinary video lowers the cost of measuring, inspecting and explaining physical places.

Evidence boundary
  • - Every figure in this note comes from a published lab experiment with its method, caveats and failures listed.
  • - Measurements from a single camera rely on an assumed scale (a storey height or a person's height) and carry the error stated with each one.
spatial intelligencespatial AI3Ddefinitions
Published by Simam Digital Ltd / Simam AI Lab Research Archive