
- 1 · Find
- 2 · Sculpture
- 3 · Before
- 4 · Measured
- 5 · True scale
Video art that turns a clip into a block of time, frames stacked like slices, has been going round this year; those pieces use brightness as fake depth and a hand-drawn camera path. Ours runs in the browser on an uploaded video, in two modes: a time sculpture and a camera path. Every moment gets real AI depth, and the path is measured from the video itself: points followed from moment to moment and lifted into 3D give where the camera went, in one to four seconds. You can type what to find and watch its path through time, and a small vision-language model says what the scene is.
What we tried
- We took apart how the popular time-volume pieces work: frames stacked along a time axis, dark pixels made see-through, brightness as depth, a manually drawn camera path.
- Every moment gets depth from ZipDepth, a 12 MB model on the graphics card. The current moment is a surface textured by the video at full resolution; the others are points, a trail of light.
- Two modes: a time sculpture (cube, stack, echo, spiral) and a camera path measured from the video. Between each pair of moments a few hundred textured points are followed, lifted into 3D with the AI depth, and a robust fit (RANSAC, then least squares) finds the camera's turn and movement, also correcting the depth's unknown offset. The moves chain into the path, which is stood on the ground found in the depth.
- We checked the measured path against COLMAP photogrammetry on four clips: our factory showcase, a drone flight, and two of our own phone clips, a courtyard walk and a walk round a car.
- To find things, YOLOS-tiny plus a colour name for every box, so typing "blue car" works; a patch matcher guided by the picture's slide and the detector's boxes follow it.
- To say what's there, SmolVLM (500M or 256M parameters) describes the start, middle and end, guesses the kind of place, and answers questions about any moment. The file's own metadata is read directly, and movement is measured: how far the camera turned, and whether a tracked object moved by itself.
- We tested on our own 42-second Quest passthrough recording of an F1 app in a front garden, using the 12 seconds from 15 s, one moment every 0.2 s.
What we measured
| Measure | First try | Shipped | Note |
|---|---|---|---|
| Camera path, courtyard walk | 5.4% | 2.7% | Position error as a share of the path, after lining up with COLMAP; before is a straight line. 100 moments, 4.1 s |
| Camera path, curving virtual camera | 11.8% | 4.4% | Our factory showcase; 51 moments, 2.1 s. Found 27 of its 37 degrees of turn |
| Camera path, straight drone flight | 2.3% | 3.3% | A straight line wins on a straight flight; ours drifts a little. 101 moments, 1.4 s |
| Camera path, circling a glossy car | 10.3% | 16% | Failed: 3 of 67 degrees of turn found. Close handheld orbits are its weak spot |
| Measuring the path | 11.6 s | 1.4 s | 101 moments; summed tables for patch matching and written-out derivatives |
| Blue car found and followed | — | 54 of 61 moments | The other 7 were whip-pans or cuts, skipped on purpose |
| Blue car, movement of its own | — | 1.8% of the frame a second | After allowing for the camera over one-second windows: correctly read as parked |
| Camera movement, measured | — | 14° left overall, 81° in all | 8 moments of whip-pan blur or a cut |
| Detector per moment | 1.3 s | 0.08 s | At its default 512-1333 px against 384 px, same objects found |
| Description per moment | nonsense | 3.1 s | SmolVLM 500M; the q4f16 weights produced nonsense, q4 works. Model 390 MB, cached after the first load |
| Current moment | 320 px dots | 720 px video | Full source resolution, bent by the depth map |
| Build 61 moments with depth | — | 6.2 s | ZipDepth 22-32 ms per moment, RTX 4070 |
What went wrong
- With full depth relief, neighbouring slices of the sculpture ran into each other and the block turned to mush. In the sculpture, depth is now a light emboss.
- The current moment looked pixelated next to the reference pieces, because it was points from a 320-pixel copy. It is now a surface textured straight from the video.
- Our first tracker locked onto a motion-smeared whip-pan and drew a confident line to nowhere. Moments that are cuts or less than half as sharp as usual are now skipped, not guessed.
- The first vision-language setup wrote nonsense ("The answer to the question is...") because of its most compressed weights; the q4 weights describe the scene properly, and the 500M model is clearly better than the 256M one.
- Our movement check first called the parked blue car "moving by itself": the camera's slide is measured in whole pixels of a small thumbnail, which is 5% of the frame a second at this frame rate. Comparing over one-second windows fixed it.
- The detector took 1.3 s a moment because it enlarges images to at least 512 px; at 384 px it finds the same objects in 80 ms. On the processor it took 29 s, so finding needs WebGPU.
- Our camera path used to be a preset shape, so a straight walk came out as an orbit. Measuring it properly worked first time on a smooth virtual camera and failed on a walk round a car: the car stays centred as you circle it, and with the depth's offset wrong, the fit explains the whole move as standing still. We checked every stage against photogrammetry: the point matches were right to 4 px, and even OpenCV's textbook method on full-size frames couldn't recover that turn from pairs 0.2 s apart. Solving for the depth offset and throwing out reflections didn't rescue it either; it is a limit of comparing neighbouring moments, and we say so on the page.
- Predicting where points would land from the previous step let one bad step derail the next. We now search around the whole picture's slide, which can't run away.
- Portrait phone video was starved: our grey working copy was 108 pixels wide. It is now sized by the short side.
What happens next
- Masks instead of boxes, with SAM 2.1 or EdgeTAM (both Apache-2.0), and anything-you-can-name search with SAM 3.1 once we've reviewed its licence.
- Asking the model about the whole clip at once, not one moment at a time.
- Longer-range matching (keyframes a second or more apart) for orbits, where neighbouring moments don't carry enough turn to measure.
- A timeline of events along the path: where the camera was when the model saw each thing, so a walk-through becomes searchable by place and time.
Built with
- ZipDepth (depth) MIT
- onnxruntime-web MIT
- YOLOS-tiny (detector) Apache-2.0
- SmolVLM-500M / 256M-Instruct (vision-language) Apache-2.0
- Transformers.js Apache-2.0
- three.js MIT
Next experiment
4D video playback: a flat video rebuilt as a moment you can walk around →
Want this for your data?