← Lab bench
EXP-057Live

4D video playback: a flat video rebuilt as a moment you can walk around

11min

from a flat 5.5-second video to a scene you can move through in space and time, on one RTX 4070

A frame of the screen-recorded factory cell with robots and workers
1/6One frame of the input: a 5.5-second screen recording of our factory cell, camera orbiting.
  1. 1 · Source
  2. 2 · Rebuilt
  3. 3 · New angle
  4. 4 · Ghost trail
  5. 5 · Spacetime
  6. 6 · Capture report

Clips of 4D reconstruction are everywhere right now: a video turned into 3D frame by frame, each frame left where and when it happened, so moving things leave trails through space. We built our own pipeline from open parts and ran it on a screen recording of our factory showcase. The still hall became a Gaussian splat; every robot and worker that moves was rebuilt in 3D for each of 155 frames. In the browser you can scrub time, run it backwards, freeze a moment and walk round it, or leave ghost trails. A capture report runs underneath: how much of each frame moves, how fast the camera turns, a plan view of where it went, events to jump to, and a build-up of every moving point in the order it was captured.

What we tried

  • COLMAP to solve the camera for every frame, then Video Depth Anything Small for steady per-frame depth, fitted to COLMAP's 3D points so every frame shares one scale.
  • A motion test with no segmentation model: each pixel is projected into frames up to two-thirds of a second away, and counts as moving if those frames see through the spot or see a different colour there.
  • LichtFeld Studio trained the still hall as a Gaussian splat, ignoring the moving pixels; the moving pixels of each frame became 3D points, 10 bytes each.
  • A browser player on Spark and three.js: the splat for the hall, one GPU buffer for all 155 frames of moving points, time scrubbing, reverse, ghost trails, a spacetime view and a ride-along 'step into the shot'.
  • A capture report, after a LiDAR replay tool we admired: per-frame telemetry, events found from the numbers, a plan view, colour by capture time or height, and a coverage box that grows as points are recorded. The share of the frame moving and the turn rate are real units; travel and the plan view are relative, because one camera has no scale.
  • Also ran a drone clip of a motorway (a third party's footage, R&D only): the pipeline worked end to end, but the cars were too small and far away to read as trails.

What we measured

MeasureStill world onlyPlus rebuilt moving pointsNote
Match to the video in moving regions14.8 dB17.3 dBPSNR from the original camera, mean of 5 frames
Match over the whole picture23.7 dB23.9 dB
Camera solved—165 of 165 frames0.7 px reprojection error
Depth error on held-out points—2.4%Median; 3.3% at the 90th percentile
Busiest moment—1.6 s1.6% of the frame moving there; 0.8% on average
Fastest camera turn—81°/sAt 1.7 s; the camera turns 38° over the clip and never holds still
Download—13 MBSplat 9.0 MB, moving points 3.5 MB, video 0.3 MB
Processing, 5.5 s clip—about 11 minCamera solve 5.5 min (CPU), depth 43 s, motion 49 s, splat 2 min 50 s

What went wrong

  • Our first motion test only asked whether later frames could see through a spot. It missed almost every robot, because most of them work in place; adding the colour check found them.
  • It also flagged the ceiling light strips and the rim of the floor against the black void outside the hall, where depth is guesswork. We now skip the void and its rim, drop hairline-shaped regions, and keep moving points within about 2 m of the floor and inside the working area.
  • The last ten frames blew up because the showcase camera speeds up just before it cuts to the next shot, so the clip stops at frame 155.
  • Our first quality check said the moving points made the match worse. The check itself was wrong: it had left the ghost-trail layer on, painting older frames over each capture.
  • Free orbit showed a hall full of fog and flat cut-out robots from the far side. That is the honest limit of one camera, so the viewer now stays near the camera's path.

What happens next

  • Our own phone footage of people moving, closer than a drone.
  • Tracking each moving thing as one object across frames, so it can be seen from more sides and trails stop shimmering.
  • Several videos placed into one scene, so different moments and seasons share a place.
  • Stereo and WebXR, to stand inside the moment in a headset.

Built with

  • COLMAP (camera solving) BSD-3-Clause
  • Video Depth Anything Small Apache-2.0
  • LichtFeld Studio (splat training, SOG export) GPL-3.0
  • OpenCV (region shapes) Apache-2.0
  • Spark 2.2 MIT
  • three.js MIT