← Lab bench
EXP-061Live

A recording you can step into, and a voice that says what happened

2.45m

closest a forklift came to a walking worker, measured from one ordinary video, with the people in it as the ruler

A runner in white rebuilt as golden-blue light with echoes, standing on a patch of beach that dissolves into darkness, the original clip inset
1/4A beach run from stock footage, projected: the runner as light with echoes of the last two seconds, the sand around him fading out. Inset: the flat original.
  1. 1 · Memory
  2. 2 · Near miss
  3. 3 · Movement review
  4. 4 · 1901

The hologram home movie in Minority Report is mostly buildable now. We take an ordinary clip, cut out every person and machine, lift them into 3D in metres, and play them back as light inside surroundings that dissolve into the dark, with the original video beside them. A small vision-language model watches along, says what it sees, guesses what happens next and answers an either-or question about the next two seconds, which the 3D tracks then mark right or wrong. Four recordings: a memory on a beach, a forklift near miss, an anonymised movement review of a busy concourse, and miners leaving Pendlebury Colliery in 1901.

What we tried

  • Video Depth Anything Small for steady per-frame depth; Grounding DINO to find people and forklifts from a text prompt; SAM 2.1 to cut each one out; a simple overlap tracker to keep identities.
  • Metres from the people: each standing person's pixel height gives their distance if they are 1.7 m tall, and their head must also sit 1.7 m above the floor the depth reveals. Searching for the one depth scale that satisfies both puts the scene in metres. On a concourse filmed from high up, over two levels, the floor comes from people's feet instead.
  • Checks that don't depend on that assumption where we could find one: the forklift, never used as a ruler, measures 1.35 m wide, against 1.1 to 1.3 m for counterbalance forklifts.
  • Cohere North Micro Vision (2.4B, Apache-2.0) running locally: every second it says what is happening, guesses what comes next, and answers one either-or question about the next two seconds (gap smaller or larger, runner closer or further, more or fewer people). The measured tracks decide right or wrong.
  • A browser player on three.js: people and machines as points of light with echoes, surroundings as a depth-bent sheet that dissolves with distance, a projector beam, floor paths, a forklift zone and gap line, a heat map, and hologram, true colour and anonymised looks.
  • Footage: Pexels clips (warehouse, beach, concourse) and a public-domain Mitchell & Kenyon film. People in the review and incident scenes are shown anonymised, and nobody is labelled a suspect or at fault.

What we measured

MeasureCheckMeasuredNote
Forklift to walking worker, closest—2.45 m at 2.6 s3.9 s inside a 3 m zone; shortest time to contact 2.3 s at the closing rate of that moment
Forklift width, never used for scale1.1–1.3 m typical1.35 mAn independent check on the metres
Rulers agreeing with the depth fit—3.2% / 3.8%Median, warehouse / beach. 6.8% on the concourse and 13.7% in the 1901 crowd, where people hide each other
Heads above the floor1.7 m assumed1.67–1.72 mFour recordings; the floor comes from the depth, the height from the boxes
Either-or calls about the next 2 scommonest answer7 of 18 rightAlways giving the commonest answer does as well or better in all four recordings (it ties in the 1901 film). The model also dodged 12 of the 39 questions
Letting the model mark its own guesses—abandonedIt approved the lazy guess "nothing changes" 9 times out of 9
Busiest moment on the concourse—21 people3 stopped for 2 s or more; median walking pace 0.43 m/s
Miners in 1901—0.70 m/s median20 walkers; the film's speed was corrected by its uploader, so treat as approximate
Per clip, RTX 4070—about 5 minDepth 0.1-1 s a frame (the GPU was shared), finding and cutting out 0.6 s a frame, narration about a minute
Download per recording—8-16 MBPoints 12 bytes each, surroundings as a 160-wide depth grid per frame

What went wrong

  • The forklift was first used as a second ruler at 2.2 m tall; its mast is raised well above that, and it came out 4.3 m. Machines are no longer rulers.
  • With only one person at one distance, size alone can't fix the depth scale: the first fit made the walker 1.31 m tall. Adding the head-height check fixed it.
  • Background seen through gaps in the forklift's mast was counted as forklift, at 80 m, making it 15 m tall. Points far from their object's median depth are now dropped.
  • On the concourse, filmed from high over two levels, the floor fit chose the wrong surface and put the camera 12 m underground. The floor now comes from people's feet.
  • The floor fit once built a 40,000-by-40,000 matrix by accident and used 17.6 GB of memory. An economy decomposition fixed it.
  • The beach camera runs with the runner, so his speed over the sand read 0.3 m/s. The page now shows only what a moving camera can measure: his distance to the lens and how fast he closes on it.
  • The small vision model rambles and repeats itself unless held to short answers, and often answers an either-or question with a word instead of the letter, or not at all.

What happens next

  • A stronger vision model (Qwen-class, 4-8B) for the voice, scored the same way against the commonest-answer baseline.
  • Construction progress: a drone pass over a housing site, every plot's build stage mapped from one flight.
  • Several cameras at once, following different subjects, played back on a real map in 3D.
  • Scrubbing time with your hand through the webcam, and standing inside the memory in a headset.

Built with

  • Video Depth Anything Small (depth) Apache-2.0
  • Grounding DINO base (finding) Apache-2.0
  • SAM 2.1 hiera-large (cutting out) Apache-2.0
  • Cohere North Micro Vision Instruct (voice) Apache-2.0
  • three.js MIT