EXP-057 · live demoGet your video in 4D →
4D Video Playback · Spatial Intelligence, Reimagined

Step inside a video, and walk around a moment.

A flat screen recording of our factory cell, rebuilt as a scene you can move through in space and in time. The still world became a Gaussian splat; every person and robot that moves was rebuilt in 3D, frame by frame, exactly where and when it was. Freeze time and walk around it, run it backwards, or leave ghost trails. Underneath, a capture report measures every frame: how much of the picture is moving, how fast the camera turns, where it went, and the moments worth jumping to.

0.00SMode · Now
Frame 1 / –
Drag to orbit · right-drag to pan · scroll to zoom
Loading the 4D scene…
Events
Speed
Time
Colour
Layers
Moving points this frame
–
Recorded so far
–
Coverage so far
–
Moving points, whole clip
–
Downloaded
–
Drawing
–

Capture record

How, where and how well this clip was captured

Source
Screen recording of our factory cellEXP-043, real-time 3D, our own footage
Clip
–
Lens
–
Camera solve
165 of 165 frames placedCOLMAP, 0.7 px reprojection error
Depth
2.4% median errorVideo Depth Anything Small, on COLMAP points held back from the fit
Camera path
–
Busiest moment
–
Camera holds still
–
Moving points
–
Still world
Gaussian splat, 9.0 MBLichtFeld Studio, moving pixels ignored in training
Match to the video
14.8 → 17.3 dBinside moving regions, from the source camera, without and with the rebuilt points
Processing
About 11 minutesone RTX 4070, mostly camera solving on the CPU

How we built it

From a flat video to a scene you can walk through in time

  1. Find the camera

    COLMAP matches features between frames and solves where the camera was for every frame: all 165, with 0.7 pixels of error. Features on moving things disagree with the rest, so the solve rests on the still world.

  2. Estimate depth

    Video Depth Anything Small gives every frame a depth map that stays steady over time. Its scale is unknown, so each frame is fitted to COLMAP's 3D points, which puts all frames in the same units. On points held back from the fit, the depth is within 2.4% (median).

  3. Find what moves

    Each pixel becomes a 3D point, projected into frames up to two-thirds of a second earlier and later. If those frames see straight through the spot where the point was, or see a different colour there, something moved. No segmentation model, so it works on anything that moves, including a robot working in place.

  4. Rebuild the still world

    LichtFeld Studio trains a Gaussian splat of the hall from the same frames, told to ignore the moving pixels, so people and robots don't leave smears in it.

  5. Rebuild every moment

    The moving pixels of each frame become 3D points in the same space, 10 bytes each: 350,000 points over 155 frames, 3.5 MB. The player keeps them all on the graphics card, so any moment, trail or tube is one draw call away. Seen from the original camera, adding them lifts the match to the video inside the moving regions from 14.8 dB to 17.3 dB.

  6. Report the capture

    The player measures the clip as it loads: the share of each frame that moves, how fast the camera turns and travels, and a growing box around everything that has moved so far. Peaks in movement, the fastest turn and any stretch where the camera holds still become events you can jump to, and a plan view shows where the camera was and what it could see. It is a record of how, where and how well the clip was captured, not only the result.

Limits

  • Everything comes from one camera, so the orbit is kept within about 15° of the path it flew. Further out, the still world turns to fog and the people and robots are revealed as painted cut-outs, because one viewpoint never saw their far side.
  • The source is a screen recording of our own real-time 3D factory, so it is clean, sharp footage. Phone video adds motion blur and rolling shutter, which make camera solving and depth harder.
  • Moving things are rebuilt as coloured points each frame, not as tracked, persistent objects. That is what makes trails easy, and also why they shimmer slightly. A few still edges are picked up as moving, and robots that stand still for the whole clip stay part of the still world.
  • Two of the report's measures are in real units: the share of the frame moving and the camera's turn rate in degrees per second. Camera travel, the plan view and the coverage box are relative, because one camera cannot tell a small room from a large one. "Moving" means the geometry disagreed between frames; it does not say what moved.
  • Processing the 5.5-second clip took about 11 minutes on one RTX 4070, most of it camera solving on the CPU. This is offline reconstruction, not live.