Location-aware apps for heritage sites: start with the place, not the pin
Most heritage apps drop a pin on a map and play a clip when you arrive. A 3D capture of the site lets the story attach to the stones, paths and views themselves, and it can come from a single drone flight.
Search for location-aware apps in the heritage industry and you find the same pattern again and again: a map, some pins, and a piece of audio or text that plays when your phone says you have arrived. It works, and visitors understand it. But it treats the site as a point on a map, when the thing people came to see is a place: a wall you can stand under, a view that opens up at a turn in the path, a building that was not always there.
Our working view is that the place itself should be the interface, an idea we first set out in place systems for heritage and destinations. Capture the site in 3D, attach the stories to the real features in that 3D, and let location decide which part of the place the visitor is shown, not just which clip to play.
What a pin can and cannot do
A geofence knows roughly where a phone is. Outdoors, consumer GPS is usually good to a few metres, and it gets worse beside tall walls, under trees and inside ruins, which is where many heritage sites are. More importantly, a pin does not know what the visitor is looking at. "You are near the north gate" is a different experience from "the gate in front of you was rebuilt in 1740, and this is where the original stood".
Content tied to pins also tends to be flat: a photo, a paragraph, an audio clip. It cannot show the site from the air, change the time period, or let someone who cannot climb the steps see what is at the top.
The place as the interface
A Gaussian splat is a photo-real 3D model built from ordinary video. It runs in a web browser, so the same capture works on the site's own website before a visit, on a visitor's phone at the gate and in a headset afterwards. We have built several of these in the lab, and three are relevant to heritage.
- Giza, rebuilt: the plateau from one 65-second stock drone clip. All 195 frames were placed by the camera solve at 0.56 px error, and the model scores 35.4 dB on frames it never saw during training. A timeline then raises Khufu, Khafre and Menkaure in the order they were built, beside AI pictures of each period, clearly labelled as illustrations (ten pictures cost $0.68).
- Our local park: two ordinary walks with a 360 camera, turned into a 3D park you can step into. A 20-second walk took 52 minutes end to end on one desktop graphics card.
- Memory Matrix: five streets in Tokyo and London, each with its recent past stacked behind it as a block of time. It is a first step towards showing a place as it was, not only as it is.
What made it work, and what did not
The honest lessons are the useful ones. At Giza the drone flew one side of the plateau, so the far faces of the pyramids are soft: capture has to plan for every side a visitor will see. Automatic detection mistook a city building for a pyramid peak, so each monument was marked by hand on one frame. And when we fitted the model to map coordinates using the three pyramids, the fit was off by 10 to 21 m. That is fine for a website, but not good enough to let a phone's GPS choose a viewpoint on site without surveyed control points.
Generated content needs a firm line, too. A picture of the Sphinx freshly painted is a strong way to tell the story, but it is an illustration, and the interface has to say so every time it appears. We wrote about that boundary in what we will not claim from a beautiful prototype.
What a heritage team needs for a location-aware 3D app
- A capture plan. A drone orbit around every face that matters, plus walking captures of the paths visitors actually use. Overcast light avoids hard shadows baked into the model.
- Geo-registration. A few surveyed points, so the model and a phone agree on where things are to within a metre or two, not tens of metres.
- Stories anchored to features. Each piece of content is tied to a 3D location and a direction to look, not just a radius on a map.
- Clear labels separating what was captured, what was reconstructed and what was imagined.
- Fast delivery. In our streaming test, the first 3D view appeared in half a second on a 20 Mbit/s connection, against 5.6 seconds to download the whole scene first. That matters on site, where signal is often poor.
Where this goes next
The open question is on-site positioning: using a phone's GPS, compass and camera to pick the right viewpoint in the captured model and hold it steady as the visitor walks. We have not tested that yet, and we will publish the result either way. If you run a site and want to try this on your own place, the lab bench shows everything we have measured so far.
Heritage sites sell their story, not only their stones. A captured 3D site gives one asset that works on the website before a visit, on a phone at the gate and in a headset later, and it can be refreshed with another flight instead of a new production.
- - Giza: a 65-second stock drone clip, 195 of 195 frames placed at 0.56 px, 35.4 dB PSNR on held-out frames, trained in 58 minutes on one RTX 4070.
- - Fitting that model to map coordinates using the three pyramids left 10 to 21 m of error, so a phone's GPS cannot yet drive the content directly.
- - We have not yet tested on-site phone positioning against a captured model. That is the open experiment.