Ordinary video, lifted into a 3D memory you can walk around. People and machines become light, measured in metres from the people in the shot. A small vision model narrates what it sees and guesses what happens next, and each guess is checked against what the 3D tracks really did.
This recording
How we built it
Video Depth Anything Small gives each frame a depth map that stays steady over time. Its scale is unknown: it knows what is nearer, not how many metres away.
Grounding DINO finds people and forklifts from a text prompt, SAM 2.1 cuts each one out to the pixel, and a simple tracker keeps the same identity from frame to frame. A person sitting inside a forklift's outline is counted as its driver.
Every standing person is used as a ruler, assumed 1.7 m tall. Their height in pixels gives their distance; their head must also sit 1.7 m above the floor that the depth reveals. The one depth scale that satisfies both puts the whole scene in metres. Machines are never used as rulers: forklift masts come in too many heights.
Every second, Cohere's North Micro Vision (2.4B parameters, running locally) looks at the last 1.5 seconds and says what is happening, then guesses what comes next. It also answers one either-or question about the next two seconds, such as "will the gap shrink or grow?", and the 3D tracks decide whether it was right. The Regent Street memory uses the larger Qwen3-VL 4B: richer descriptions, and scored the same way.
For busy streets the memory is split into chapters where the number of people changes level, the ground lights up wherever people walked, and each walker's next two seconds are guessed by keeping them going straight at their recent pace. The recording itself says where they really went, so every guess is scored, next to the lazy guess that everyone stands still.
Freeze any moment and each walker splits into three futures for the next two seconds: keep going straight at their pace, slow to a stop, or follow the crowd (steer the way everyone else moved through that patch of pavement, with the walker's own steps left out). We tried to pick the likeliest in advance, from which rule best explained the walker's previous second, and it did worse than a blind pick, so the three are drawn as equals and you place the bet. Press play and the recording shows which branch came true; every freeze in the clip is scored the same way.
On asphalt after dark, a spot much brighter than the road around it is a light mirrored in water, so the share of the visible road doing that is measured frame by frame. Grounding DINO looks for umbrellas once a second. Two gates are drawn across the pavement and the road, and every tracked path that crosses one is counted, with its direction and its speed at the line.
In the browser, the people and machines of each frame are drawn as points of light where they really stood, the surroundings dissolve around them, and echoes show where they have been. The measurements are drawn into the scene.
Limits