EXP-061 · live demoGet this for your venue →
Memory Projector · video to 4D, with a voice

Step into a recording, and ask it what happened.

Ordinary video, lifted into a 3D memory you can walk around. People and machines become light, measured in metres from the people in the shot. A small vision model narrates what it sees and guesses what happens next, and each guess is checked against what the 3D tracks really did.

0.00SLook · Hologram
–
Original video
Loading the memory…
Look
Layers
Speed

This recording

What the memory measured

How we built it

From a flat recording to a memory with a voice

  1. Depth for every frame

    Video Depth Anything Small gives each frame a depth map that stays steady over time. Its scale is unknown: it knows what is nearer, not how many metres away.

  2. Find and cut out who and what is there

    Grounding DINO finds people and forklifts from a text prompt, SAM 2.1 cuts each one out to the pixel, and a simple tracker keeps the same identity from frame to frame. A person sitting inside a forklift's outline is counted as its driver.

  3. Metres from the people themselves

    Every standing person is used as a ruler, assumed 1.7 m tall. Their height in pixels gives their distance; their head must also sit 1.7 m above the floor that the depth reveals. The one depth scale that satisfies both puts the whole scene in metres. Machines are never used as rulers: forklift masts come in too many heights.

  4. A voice that sees, guesses, and is checked

    Every second, Cohere's North Micro Vision (2.4B parameters, running locally) looks at the last 1.5 seconds and says what is happening, then guesses what comes next. It also answers one either-or question about the next two seconds, such as "will the gap shrink or grow?", and the 3D tracks decide whether it was right. The Regent Street memory uses the larger Qwen3-VL 4B: richer descriptions, and scored the same way.

  5. Chapters, flow and prediction

    For busy streets the memory is split into chapters where the number of people changes level, the ground lights up wherever people walked, and each walker's next two seconds are guessed by keeping them going straight at their recent pace. The recording itself says where they really went, so every guess is scored, next to the lazy guess that everyone stands still.

  6. Possible futures

    Freeze any moment and each walker splits into three futures for the next two seconds: keep going straight at their pace, slow to a stop, or follow the crowd (steer the way everyone else moved through that patch of pavement, with the walker's own steps left out). We tried to pick the likeliest in advance, from which rule best explained the walker's previous second, and it did worse than a blind pick, so the three are drawn as equals and you place the bet. Press play and the recording shows which branch came true; every freeze in the clip is scored the same way.

  7. Street conditions

    On asphalt after dark, a spot much brighter than the road around it is a light mirrored in water, so the share of the visible road doing that is measured frame by frame. Grounding DINO looks for umbrellas once a second. Two gates are drawn across the pavement and the road, and every tracked path that crosses one is counted, with its direction and its speed at the line.

  8. Project it

    In the browser, the people and machines of each frame are drawn as points of light where they really stood, the surroundings dissolve around them, and echoes show where they have been. The measurements are drawn into the scene.

Limits

  • One camera sees one side of everything: move far off the camera's line and people become flat cut-outs. The orbit is limited for that reason.
  • Metres rest on the 1.7 m assumption. Real people vary by about ±7%, so single distances are good to roughly that; the page shows how well the rulers agreed with each other and an independent check where one exists.
  • The vision model is small. Its sentences are often plain and repetitive, and its free-text guesses are not scored. Only the either-or answers are scored, next to the score you would get by always giving the most common answer. We first let it mark its own guesses: it approved the lazy guess "nothing changes" nine times out of nine, so we stopped.
  • On busy streets people are 6 to 30 metres away, and depth that far out jitters by tens of centimetres per frame. Paths are smoothed over a second; close passes between people and vehicles are not shown, because people standing behind parked cars merge with them in depth.
  • The wet-road figure is the share of the road that mirrors a light, not millimetres of rain: a camera can't measure rainfall, and it isn't calibrated against the same street dry. With the place and time of a recording, weather records could sit alongside it; stock footage doesn't come with them.
  • Vehicle speeds rest on the same people-based scale. The buses measure 4.0 m tall against a real 4.4 m, so speeds may read up to 10% low. Only vehicles within 30 m get a speed tag; further out the depth is too rough.
  • The futures are three simple rules, not a model trained on people, and two seconds is a short horizon. They are scored against the same recording they are drawn on, next to the lazy answer of always picking the rule that wins most; where they don't beat it, the page says so.
  • This is movement review, not identification. The Anonymised look removes appearance entirely, nobody is labelled a suspect or at fault, and no face recognition is used anywhere.