Searching inside a 3D scene: how AI finds “trees” or “the rail line” in a Gaussian splat
A 3D splat is a million coloured blobs with no labels. We made one searchable by running open AI models on the video frames and lifting their answers into 3D, without retraining the model.
A Gaussian splat looks like a photograph you can walk around, but underneath it is about a million small coloured blobs with no idea of what they belong to. Nothing in the file says "tree" or "road". To make our apartment viewer searchable, we had to give those blobs meaning, and we wanted to do it without retraining the 3D model every time someone thinks of a new word.
Ask the frames, then lift the answers into 3D
Image AI is very good at finding things in photographs, and the 3D model was built from photographs: the frames of the drone clip. So the search runs on the frames, and the results are carried into 3D.
- Find: for each search word, Grounding DINO, an open-vocabulary detector, draws boxes around matches in 44 of the frames.
- Cut out: SAM 2.1 turns each box into a precise mask.
- Lift: every point of the 3D model is projected into each frame that sees it. Its score is the average confidence over those frames, and it only counts if at least three frames saw it. One stray detection cannot light up a point on its own.
- Ship: the scores go to the browser as one byte per point per word, 1.8 MB for 150,312 points and six words, alongside the 14.4 MB model.
In the viewer, a search shades every match from pink to white by confidence, and adds up the matching area in 25 m squares. The guide figures for the site are about 3,900 m² of trees, 3,700 m² of green space, 1,300 m² of rail line and 1,000 m² of road.
Small things, big things
Windows were too small to find in a whole drone frame, so the search for them runs on tiles of each frame. Cutting out 150 windows at once filled the graphics card's memory, so small things keep their boxes as the mask. Areas like lawns and roads are the opposite: their boxes are huge, and a filter meant to reject oversized boxes was throwing them away until we let area words through.
The words that failed
We tried ten words and kept six: trees, green space, windows, roads, rail line and cars. "Footpath" took the whole lawn. "Balcony", "rooftop" and "crane" were dropped after we checked their masks against the frames. An early "rail" search found the viaduct instead of the tracks until the prompt was reworded. Dropping a word that does not work is better than shipping a confident wrong answer.
We also chose not to offer a search for people. A site viewer has no reason to find individuals in footage, and we do not put identity information on strangers.
Why not train language into the splat?
Research methods such as LangSplat train language features into every Gaussian, which allows truly free-text queries. They cost a training run per scene and a heavier download. Our approach needs no retraining: adding a word means running two open models on the frames again, which takes minutes on one graphics card. For a site with a known vocabulary (trees, roads, windows, parking) that trade-off suits us. Free-text search is on our list.
What comes next
- Score the search against hand-labelled frames before quoting any accuracy.
- Free-text queries that match any words, not a fixed list.
- The same search on our own drone flights of housing sites, where "plots with foundations" is the question that matters.
This is one part of what we mean by spatial intelligence: a model of a place you can ask questions of, not just look at.
Search turns a 3D capture from something to look at into something to query: how much green space, where the rail line runs, which facades have windows. It is the first step from a pretty model to a site you can ask questions of.
- - Six of ten search words were kept after checking their results by eye. The search has not yet been scored against hand-labelled frames, so we quote no accuracy figure.
- - Area totals depend on the opacity threshold and grid size used, and are a guide, not a survey.