A flat screen recording of our factory cell, rebuilt as a scene you can move through in space and in time. The still world became a Gaussian splat; every person and robot that moves was rebuilt in 3D, frame by frame, exactly where and when it was. Freeze time and walk around it, run it backwards, or leave ghost trails. Underneath, a capture report measures every frame: how much of the picture is moving, how fast the camera turns, where it went, and the moments worth jumping to.
Capture record
How we built it
COLMAP matches features between frames and solves where the camera was for every frame: all 165, with 0.7 pixels of error. Features on moving things disagree with the rest, so the solve rests on the still world.
Video Depth Anything Small gives every frame a depth map that stays steady over time. Its scale is unknown, so each frame is fitted to COLMAP's 3D points, which puts all frames in the same units. On points held back from the fit, the depth is within 2.4% (median).
Each pixel becomes a 3D point, projected into frames up to two-thirds of a second earlier and later. If those frames see straight through the spot where the point was, or see a different colour there, something moved. No segmentation model, so it works on anything that moves, including a robot working in place.
LichtFeld Studio trains a Gaussian splat of the hall from the same frames, told to ignore the moving pixels, so people and robots don't leave smears in it.
The moving pixels of each frame become 3D points in the same space, 10 bytes each: 350,000 points over 155 frames, 3.5 MB. The player keeps them all on the graphics card, so any moment, trail or tube is one draw call away. Seen from the original camera, adding them lifts the match to the video inside the moving regions from 14.8 dB to 17.3 dB.
The player measures the clip as it loads: the share of each frame that moves, how fast the camera turns and travels, and a growing box around everything that has moved so far. Peaks in movement, the fastest turn and any stretch where the camera holds still become events you can jump to, and a plan view shows where the camera was and what it could see. It is a record of how, where and how well the clip was captured, not only the result.
Limits