
- 1 · The site
- 2 · Stages
- 3 · Sequence
- 4 · Plots
Drone progress reports are a real business in construction. We tried the core of one on free footage of a UK-style new-build site: the flight is solved and rebuilt as a Gaussian splat you can fly around, an open-vocabulary detector finds every plot in every frame, and detections that land on the same spot become one plot with a pin. Colour shows how far each plot has got: foundations, slab, structure going up, or roofed. Clicking a pin shows the frame it was seen best in. A vision model running on the same machine, Qwen3-VL-4B, reads each plot's stage, and we score it against labels made by eye. A site report adds what a site manager would ask: machines and materials counted, metres from the parked cars, a walkable / caution / blocked map of the ground with the walk from the site cabin to every unfinished plot, and a short briefing written by the AI, each figure tagged as measured, estimated or the AI's guess.
What we tried
- 131 frames, three a second, from a Pexels drone clip. COLMAP features and sequential matching, GLOMAP for the solve: all 131 frames placed, 82,582 points.
- LichtFeld Studio trained a Gaussian splat of the site: 15,000 steps, 1.5 million splats, 5 min 48 s on one RTX 4070, 21 MB compressed.
- Grounding DINO searched every frame twice, once for houses and once for slabs, foundations and frames: 3,152 boxes. Each box is placed in 3D by the solved points inside it; boxes from different frames within half a house-width become one plot, kept if seen in three frames or more: 120 places, 113 of them real plots.
- Every plot's best view was labelled by eye into foundations, slab, structure or roofed, as a yardstick.
- Stage reading, attempt 1: Cohere North Micro Vision (2.4B) choosing one of seven stages. It said "slab" for 117 of 120.
- Attempt 2: the same model, four yes/no questions per plot (a plot? roofed? walls? slab?). It called most real houses "not a plot": 27 of 120 right.
- Attempt 3: Grounding DINO scoring stage phrases (tiled roof, scaffolding, concrete slab, foundation trench) on each crop, with thresholds fitted on half the plots and tested on the other half: 68% on the held-out half, exactly what "always roofed" scores.
- Attempt 4: Qwen3-VL-4B (Apache-2.0, 8.9 GB, local) choosing one of five stages. It read the unfinished plots well, 25 of 29, but called 77 of 84 finished houses "structure": 27% overall.
- Attempt 5, rule fixed before running: the same model is first asked only whether the roof is finished and tiled. Yes means roofed; otherwise its five-way answer stands. Nothing was fitted to the labels, so the score over all 120 plots is fair: 81%.
- Site report: COLMAP multi-view stereo gives 1.36 million dense points and a height map of 0.9 m squares. Metres come from five parked cars taken as 4.4 m long (1 unit = 18.1 m, cars disagree by about 20%); as a check, roofed houses then stand 6.4 m tall.
- Grounding DINO with one prompt per kind of thing found machines, material stacks, vehicles, site cabins and spoil heaps; every one was checked by eye and only the confirmed ones are used.
- Qwen3-VL-4B read the ground in 921 squares of about 6 m, each cropped from the sharpest drone frame. Fixed rules (more than 1.2 m above the ground, a 1 m drop, the surface, 5 m around a machine) make a walkable / caution / blocked map, and A* finds the walk from the site cabin to each unfinished plot.
- Qwen3-VL also writes a five-point briefing for a site manager from six frames and the measured counts. It is shown as written and labelled as the AI's guess.
- Gemini Vision on one real frame: Gemini 3 Pro Image imagines the same view before any building and once the estate is finished (a drag-to-compare slider), and Gemini 3.8 Flash returns a structured read (phase, machines, material areas, hazards, lorry and walking routes, next steps) drawn over the photo.
What we measured
| Measure | Baseline | Result | Note |
|---|---|---|---|
| Frames placed by the camera solve | — | 131 of 131 | 82,582 points; 72° lens solved |
| Site rebuild | — | 5 min 48 s | 1.5M splats, 21 MB for the browser |
| Detections to plots | — | 3,152 boxes to 120 places | 113 real plots; 7 were cabins, material stacks, a sign and a mound |
| Plots by stage (by eye) | — | 84 roofed, 9 structure, 12 slab, 8 foundations | 74% roofed |
| Stage by a 2.4B vision model, 7-way | 84 of 120 (always roofed) | answered slab 117 times | Unusable |
| Stage by the same model, yes/no | 84 of 120 | 27 of 120 | |
| Stage by detector phrases, held-out half | 68% (always roofed) | 68% | No better than the baseline |
| Stage by Qwen3-VL-4B, 5-way | 70% (always roofed) | 27% | Called 77 of 84 roofed houses structure |
| Stage by Qwen3-VL-4B, roof first, all 120 | 70% (always roofed) | 81% | Held-out half: 77% against 68% |
| Unfinished plots read right (roof first) | 0 of 29 (always roofed) | 25 of 29 | Non-plots: 0 of 7 recognised |
| Site machines found, checked by eye | — | 15 of 19 real | Vans and material stacks were the misses |
| Material stacks found, checked by eye | — | 12 of 23 real | Trenches and foundations taken for stacks |
| Site cabins found, checked by eye | — | 3 of 84 real | Failed: houses and slabs called cabins; not used |
| Ground surface by Qwen3-VL, 60 random squares | 32 of 60 (all soil) | 41 of 60 | Never called a roof a building |
| Walk map (walkable / caution / blocked), same squares | 36 of 60 (all caution) | 46 of 60 | Eye labels made after seeing the model's answers |
| Unfinished plots reachable from the site cabin | — | 22 of 29 | Typical walk 200 m for 169 m as the crow flies |
What went wrong
- One detector prompt for every stage mixed them up, and boxes around whole streets of trenches swallowed several plots. Two prompts and a size limit fixed most of it; a few foundation plots are still found as one.
- The small vision model can describe a beach, but not an aerial building site: every framing of the stage question failed, and yes/no questions made it worse. A stronger model, Qwen3-VL-4B, got there only once the roof question was asked on its own.
- The floor fit accepted every point, because the tolerance scaled with a site 20 units long. The site is flat, so the floor still came out right, but the check is weaker than it looks.
- The first ground map, from the splat's own Gaussians, was 12% filled: most of a splat is soft blobs. Dense stereo fixed it. The first walk rules then turned every kerb red (noise in the dense points) and could not tell pale roads from pale bare earth by colour, so the surface reading moved to the vision model.
- The "site cabin" and "spoil heap" prompts found houses, slabs and garden beds. They are reported as failures and left off the map.
- Gemini's "safer walking route" runs right beside its own lorry route, which is what a site should keep apart. Its boxes were all on real things; its routing judgement was not.
What happens next
- Show the model two or three views of each plot, and a closer crop, to fix the far-off roofs it still misreads.
- Repeat flights of the same site, aligned into one model, for a real timelapse of progress.
Built with
- COLMAP + GLOMAP (camera solve) BSD-3-Clause
- LichtFeld Studio (splat training) GPL-3.0
- Grounding DINO base (finding plots) Apache-2.0
- Cohere North Micro Vision Instruct (stage attempts) Apache-2.0
- Qwen3-VL-4B-Instruct (stage reading, ground surface, briefing) Apache-2.0
- COLMAP multi-view stereo (dense ground) BSD-3-Clause
- Gemini 3 Pro Image and Gemini 3.8 Flash (imagined views, site read) Google Gemini API
- Spark 2.2 MIT
- three.js MIT