Every moment of your clip becomes a slice with its own AI depth. Stack the slices into a block of time, or lay them along the path the camera travelled. Type what you're looking for, "blue car" say, and a detector finds it; a glowing line then follows it through the whole sculpture. Nothing is uploaded.
How it works
The browser steps through your clip and keeps a frame every fraction of a second. Nothing leaves your device.
ZipDepth, a 12 MB depth model, estimates how far away every pixel is, on your graphics card. Each moment becomes a sheet bent by its own relief rather than a flat picture: the current one drawn from the video at full resolution, the rest as a see-through trail. Brightness-as-depth is there too, for a purely artistic look.
Between each pair of moments, a few hundred textured points are followed from one to the next and lifted into 3D with the AI depth. A robust fit finds the turn and movement of the camera that explains where they landed, also correcting the depth's unknown offset, and the moves chain into the path the camera flew. It is stood on the ground found in the depth and takes one to four seconds.
As a sculpture, time becomes the third axis: a cube, a fanned stack or an echo tunnel, optionally twisted into a spiral. Along the camera path, each moment is placed where the camera was and faces the way it looked; at true scale they line up into one 3D scene, stretched they become a trail.
YOLOS-tiny, a 26 MB detector that knows 80 everyday kinds of thing, looks at every moment in about 0.1 s each on the graphics card. We name the colour inside each box, so "blue car" or "person" finds the right one. Or click anything yourself.
SmolVLM, a 500-million-parameter vision-language model (or its 256-million sibling), describes the start, middle and end of the clip and answers questions about any moment, on your graphics card. The file's own metadata (date, device, location, if the camera stored it) is read directly. How the camera moved, and whether a tracked object moved by itself or only with the camera, is measured from the picture, not guessed by the model.
The picture's overall slide between moments predicts where the object went, a patch match finds it, and the detector snaps a real box around it. Smeared whip-pans and cuts are skipped rather than guessed. Its positions join into a glowing path through the sculpture.
Limits