
- 1 · Site
- 2 · Unit
- 3 · Trees
- 4 · Rail
We took a 29-second stock drone clip of an apartment block beside a park and timed every step to a 3D Gaussian splat: 9 minutes 37 seconds on one desktop GPU, nothing sent to the cloud. On top sits an apartment finder we built for it, Simam Spatial Residences. Heights, floors and footprints are read from the 3D. Each tower shows its sample units on the facade, coloured by availability, with filters for beds, outlook and budget. A unit opens a listing with an AI picture of its interior (the view in its windows is the real drone footage), a floor plan, and where it sits on the 3D. Around here shows the park and places nearby, and an AI search shades the scene by what it is (trees, green space, the rail line, windows) with area totals per 25 m square.
What we tried
- Kept the sharpest frame in every third of a second from a 29-second, 1080p drone clip (88 frames), solved the cameras in COLMAP and trained a splat in LichtFeld Studio, timing each step.
- Compared a short 7,500-step training run with the usual 30,000 steps on held-out frames, and kept the short one.
- Measured each building by marking its corners on drone frames and reading their positions from the splat, then drew the boxes back onto the frames to check them.
- Ran Grounding DINO and SAM 2.1 for each search word on 44 frames and averaged their confidence onto every point of the 3D. A local Qwen3-VL-4B drafted selling points, and we kept only the ones we could see.
- Asked Gemini for eight sample interiors and amenities, each started from a real drone crop of the park or skyline so the windows show what the drone saw, under a $0.60 cap.
What we measured
| Measure | Full training | Short training (used) | Note |
|---|---|---|---|
| Clip | — | 29 s, 1080p, Pexels | |
| Frames placed | — | 88 of 88 | 0.48 px camera solve error |
| Frames and camera solve | — | 3 min 58 s | |
| Training | 26 min 20 s | 5 min 39 s | 30,000 against 7,500 steps, 1M Gaussians cap |
| Clip to finished 3D | 30 min 18 s | 9 min 37 s | |
| Held-out PSNR | 36.1 dB | 34.7 dB | Every 8th frame held out; by eye the two look alike |
| Tower height | — | 34 m, 10 floors | ±12%; the neighbour's roof is 0.79 of the tower's, matching 8 floors to 10 |
| AI searches kept | — | 6 of 10 | Footpath, balcony, rooftop and crane dropped after checking the masks |
| AI interior and amenity pictures | — | 8 for $0.55 | Gemini, about $0.068 each, against a $0.60 cap |
What went wrong
- Training first failed to start: the camera model has lens distortion, which LichtFeld needs its --gut mode for.
- The full training run took 26 minutes once it reached a million Gaussians, which broke our under-20-minutes aim. A quarter of the steps came in at 5 min 39 s and looked the same on held-out frames.
- Windows were too small to find in a whole frame, and cutting out 150 of them at once filled the GPU's memory. Searching in tiles and keeping the boxes fixed both.
- A view from the windows looked into blur, because the drone never filmed outwards from the tower, and a straight-down view showed sky-coloured smears. The unit's outlook is now a wedge on the 3D plus an AI interior, and the camera stays on the side the drone flew.
- "Footpath" took the whole lawn, so it was dropped. The coal train looked like a good ruler for scale, but it came out too faint in the 3D to measure, so scale comes from counted storeys instead.
What happens next
- Score the AI search against hand-labelled frames before quoting any accuracy.
- Free-text search, matching any words rather than a fixed list.
- The same viewer on our own drone flights, with real listings from a developer.
Built with
- COLMAP 4.1 BSD-3-Clause
- LichtFeld Studio GPL-3.0
- Grounding DINO Apache 2.0
- SAM 2.1 Apache 2.0
- Qwen3-VL-4B Apache 2.0
- Gemini image model (interiors) Google API terms
- Spark 2.2 MIT
- three.js MIT