
- 1 · All five
- 2 · Memory
- 3 · Back in time
- 4 · Deck of moments
- 5 · Camera's eye
Our Memory Projector meets our Time Sculptor. Five ordinary recordings, Shibuya Crossing and four London streets, each become a 3D memory: the street solid in its own colours, the people walking through it as light. Step back and every half second of the clip stacks up behind the present as a sheet shaped like the street, with each walker's outline on it and a line of light tracing their path back through time. Stand in the middle and all five surround you, each seen from where its camera stood. Google's Gemini read each scene; we checked every fact by hand.
What we tried
- A first version used five London clips with crowds mostly 6 to 30 m away on streets running 80 m back, shown as five small cells in a ring. It looked poor: thin figures, streaked streets. We started again from footage chosen for close people.
- A person detector ranked candidate clips by how tall the nearest people stand in the frame and how much the camera drifts. The five chosen have their tallest person at 23% to 99% of the frame height; the Tokyo crossing drifted up to 18 pixels and was steadied onto its first frame.
- Video Depth Anything Small for steady depth per frame; Grounding DINO and SAM 2.1 to find and cut out every person; people as rulers (1.7 m tall, heads 1.7 m above the floor) to put each street in metres.
- The empty street: each pixel's median over the frames where nobody covers it, with Depth Anything V2 run once more on that still at a higher resolution and its detail fitted onto the video depth. Anyone who stood still long enough to stay in it is blurred and greyed, plus a few spots checked by eye.
- The block of time: every half second as a sheet shaped like the street, carrying that moment's walkers as light and outline, blended back to front as in our Time Sculptor, and pushed back along the camera's line of sight by its age.
- Google's Gemini read five frames of each memory (place, light, weather, three log lines), told never to describe a person. We checked every place and fact, rewrote one line and dropped one landmark.
What we measured
| Measure | Check | Measured | Note |
|---|---|---|---|
| People tracked | — | 260 | 40 in Tokyo, 74, 47, 47 and 52 in London; someone hidden and seen again can count twice. The busiest moment has 23 people in view, at Shibuya |
| Nearest crowd | — | 6.7 m | Half the walkers on Regent Street passed within 6.7 m of the camera; 7.8 to 11.1 m in the other four |
| Rulers agreeing with the depth fit | — | 11-13% | Median, four London streets. 24% at Shibuya, where people are seen from above |
| Heads above the floor | 1.7 m assumed | 1.62-1.69 m | All five; the floor comes from the depth, the height from the boxes |
| Still-image depth against video depth | — | r = 0.91-1.00 | How well the sharper street depth fits the video's, before its detail is added |
| Street blurred for people standing still | — | 3-5% | Three of the five; 14% on Oxford Street and 27% on Regent Street, where people stood about, and spots we checked by eye |
| Scene reading | — | under 4 US cents | Five Gemini readings, five frames each |
| Per clip, RTX 4070 | — | about 3.5 min | Depth, cut-outs, metres and export for 12 s at 12 fps, once the libraries were loaded |
| Download per memory | — | 14-18 MB | About 10,000 points of light a frame, the street, and 24 moments of the past |
What went wrong
- In the first version most people were 6 to 30 m away and small in the frame, so they read as blobs, and the streets behind them ran 80 m back, where depth from one camera swings by metres. Better footage fixed it, not better code.
- Our first test for people standing still blurred anyone covered in a third of the frames, which on a busy pavement was 25 to 40% of the street. Only spots covered in nearly every frame can keep a person, so the test now asks for 75%.
- Every sheet's edges sampled depth from the neighbouring moment in the shared image, stretching them into huge flat fans. Each lookup now stays inside its own moment.
- Twenty-four sheets of light added together burned out to white. They are now blended back to front, as the Time Sculptor's are, with walkers as outlines.
- At first the hologram people vanished: they stood beyond the distance where we flattened the street into a backdrop. The backdrop now always sits beyond the walkers.
- Seen from the side, a street lifted from one camera smears. We tried five memories in a row and from above; what worked was putting all five cameras at one point, each facing its own way.
- Gemini wrote "Night falls" for a clip filmed entirely after dark, and gave a street address we couldn't confirm. Both were corrected by hand.
What happens next
- Several cameras on the same street, so a memory can be walked round, not only looked into.
- The same street on different days, played one after another in the same place.
- A version you stand inside in a headset.
Built with
- Video Depth Anything Small (depth) Apache-2.0
- Depth Anything V2 Small (street detail) Apache-2.0
- Grounding DINO base (finding) Apache-2.0
- SAM 2.1 hiera-large (cutting out) Apache-2.0
- Google Gemini API (scene reading) Google API terms
- Pexels stock footage Pexels licence
- three.js MIT