A 44-second drone pass over a new-build housing site, rebuilt as a 3D model you can fly around. Every plot the detector found is pinned where it stands and coloured by how far it has got: foundations, slab, structure going up, or roofed.
This flight
Site report
Every figure says what it is: counted from the flight, worked out from an assumption you can change, or the AI's own opinion.
Gemini Vision
One real frame from the flight. Gemini 3 Pro Image imagines the same view before any building and once the estate is finished; Gemini 3.8 Flash reads what is going on now. Drag the line to compare.
How we built it
131 frames, three a second. COLMAP matches features between neighbouring frames and GLOMAP solves where the drone was for every one: all 131 placed, 82,582 points on the ground and buildings.
LichtFeld Studio trains a Gaussian splat from the same frames: 1.5 million splats in under six minutes on one RTX 4070, 21 MB compressed for the browser.
Grounding DINO searches each frame twice, for houses and for slabs, foundations and frames. Each box is placed in 3D by the solved points inside it, and boxes from different frames that land on the same spot become one plot: 3,152 boxes, 120 places, each seen in at least three frames.
Qwen3-VL-4B, a vision model running on the same machine, looks at each plot's best photo. Asked to pick one of five stages it called nearly every finished house "structure", so it is now asked one plain question first, is the roof finished and tiled, and only picks a stage when the answer is no. That reading matches the eye labels on 81% of plots (77% of the held-out half), where calling everything "roofed" scores 70% (68%). Three smaller models before it did no better than that shortcut. Eye labels are shown by default; switch to "AI model" to see its reading.
COLMAP multi-view stereo turns the frames into 1.36 million dense points, binned into a height map of roughly 0.9 m squares. Parked cars set the scale, and roofed houses then stand 6.4 m tall, as they should. Grounding DINO finds machines, material stacks and vehicles, each checked by eye; Qwen3-VL reads the ground in 921 squares of about 6 m. Fixed rules turn height, drops, surface and a 5 m zone around each machine into walkable, caution or blocked, and routes are found across that map.
One real frame goes to Google's Gemini API: Gemini 3 Pro Image imagines the view before building and when finished, and Gemini 3.8 Flash returns a structured read of the site with boxes and routes in image coordinates. Every box was checked by eye against the photo; the imagined views are shown as imaginings, not plans.
Limits