Five ordinary recordings, one from Tokyo and four from London, each lifted into a 3D memory measured in metres from the people in it. The street stands in its own colours; the people walk through it as light. Step back and every half second of the past stacks up behind the present as a block of time, each walker leaving a line of light through it.
Side by side
How we built it
We ranked candidate clips with a person detector: how tall the nearest people stand in the frame, and how much the camera drifts. These five have people within a few metres, filling up to the whole frame height. The Tokyo crossing drifted by up to 18 pixels, so every frame was shifted back onto the first.
Video Depth Anything Small gives each frame a depth map that stays steady over time. For the empty street, Depth Anything V2 runs once more on a still at a higher resolution, and its fine detail is fitted onto the video depth, so edges sharpen without changing the metres.
Grounding DINO finds people, cars and buses, SAM 2.1 cuts each one out to the pixel, and a tracker keeps identities from frame to frame. Each pixel of the empty street is the median of the frames where nobody covers it. Anyone who stood still long enough to survive that is blurred and greyed, along with a few spots we checked by eye.
Every standing person is a ruler, assumed 1.7 m tall: their height in pixels gives their distance, and their head must sit 1.7 m above the floor the depth reveals. The one scale that satisfies both puts the street in metres.
The empty street becomes a dense mesh in its own colours, torn where depth jumps so nothing stretches, with a smoothed copy behind it to fill the tears. The people are points of light, about 10,000 a frame, coloured by brightness only.
Every half second becomes a sheet shaped like the street, carrying that moment's walkers as light and outline. Older sheets sit further back along the camera's line of sight, so the clip stacks into a block behind the present, as in our Time Sculptor; each walker's path runs back through it as a line.
Google's Gemini looked at five frames of each memory and said where it is, the light and weather, and a few log lines about what the memory holds. It was told never to describe a person. We checked every place and fact by hand, rewrote one log line and dropped a landmark address we couldn't confirm. The readings cost under ten US cents.
Limits