EXP-058 · live demo
Video Time Sculptor · runs on your device

Turn any video into a sculpture of time.

Every moment of your clip becomes a slice with its own AI depth. Stack the slices into a block of time, or lay them along the path the camera travelled. Type what you're looking for, "blue car" say, and a detector finds it; a glowing line then follows it through the whole sculpture. Nothing is uploaded.

–
0.0Time sculpture · Cube–
Drag to orbit · right-drag to pan · scroll to zoom
Click the thing to track
Loading the sample clip…
00.0 / 00.0
Captures
–
Depth per capture
–
Points drawn
–
Finding objects
–
Describing
–
Drawing
–

How it works

A video as a block of time you can hold

  1. Capture moments

    The browser steps through your clip and keeps a frame every fraction of a second. Nothing leaves your device.

  2. Give each one depth

    ZipDepth, a 12 MB depth model, estimates how far away every pixel is, on your graphics card. Each moment becomes a sheet bent by its own relief rather than a flat picture: the current one drawn from the video at full resolution, the rest as a see-through trail. Brightness-as-depth is there too, for a purely artistic look.

  3. Measure where the camera went

    Between each pair of moments, a few hundred textured points are followed from one to the next and lifted into 3D with the AI depth. A robust fit finds the turn and movement of the camera that explains where they landed, also correcting the depth's unknown offset, and the moves chain into the path the camera flew. It is stood on the ground found in the depth and takes one to four seconds.

  4. Arrange them in time

    As a sculpture, time becomes the third axis: a cube, a fanned stack or an echo tunnel, optionally twisted into a spiral. Along the camera path, each moment is placed where the camera was and faces the way it looked; at true scale they line up into one 3D scene, stretched they become a trail.

  5. Find it

    YOLOS-tiny, a 26 MB detector that knows 80 everyday kinds of thing, looks at every moment in about 0.1 s each on the graphics card. We name the colour inside each box, so "blue car" or "person" finds the right one. Or click anything yourself.

  6. Say what it sees

    SmolVLM, a 500-million-parameter vision-language model (or its 256-million sibling), describes the start, middle and end of the clip and answers questions about any moment, on your graphics card. The file's own metadata (date, device, location, if the camera stored it) is read directly. How the camera moved, and whether a tracked object moved by itself or only with the camera, is measured from the picture, not guessed by the model.

  7. Follow it through time

    The picture's overall slide between moments predicts where the object went, a patch match finds it, and the detector snaps a real box around it. Smeared whip-pans and cuts are skipped rather than guessed. Its positions join into a glowing path through the sculpture.

Limits

  • The measured path came within 3 to 4% of a full photogrammetry solve (COLMAP) on a walk, a drone flight and a moving virtual camera, but it failed on a slow handheld circle round a glossy car, where it found 3° of a 67° turn. Close orbits and reflective subjects are its weak spot; the Orbit shape is there for those. For the offline, fully solved version, see our 4D playback demo.
  • The depth is relative: each moment's relief is right in itself, but its scale is only kept steady from one moment to the next, not measured.
  • The detector only knows 80 everyday kinds of thing (people, vehicles, animals, furniture, food), and it names colours roughly: a render of a racing car in our test was called a "grey car". Anything else, you click.
  • The vision-language model is small. Its descriptions are usually right in outline and sometimes wrong in detail, and a guess at the place is only a guess: we label it as one.
  • Tracking follows appearance. Fast whip-pans, motion blur and things leaving the frame lose it; the path then skips those moments.