EXP-052 · live demo

Flat video in, 3D video out, live.

An AI depth model runs on your own graphics card inside this page. For every frame it estimates how far away each pixel is, then a shader moves pixels sideways by their depth to build a separate left and right eye. It's converting a sample clip now. Try your webcam or your own video.

Input · flat 2D clip
Output · red/cyan 3D
Loading the depth model…
Show the output as
Depth model
Input
Depth per frame
–
3D output rate
–
Runs on
–
Depth input
–

Nothing leaves your device. The model downloads once (ZipDepth from this site, Depth Anything from Hugging Face) and runs locally; camera frames and your videos are never uploaded. Red/cyan glasses show the 3D view; side by side works in a VR headset's browser or media player.

How we built it

The same pipeline as our live server, shrunk into a web page

  1. Depth for every frame

    Two models to compare. ZipDepth (6.1M parameters, MIT, ECCV 2026), which we exported to a 12 MB half-precision ONNX file, runs through ONNX Runtime Web at 448×256. Depth Anything V2 Small (Apache-2.0) runs through Transformers.js at 476×266. Both run on WebGPU and predict relative depth: nearer or further, not metres.

  2. Keep depth steady over time

    Each frame's near and far limits are smoothed over time so the scene doesn't pulse. When a hard cut is detected (the picture changes too much in one frame), the smoothing resets so one shot's depth never leaks into the next.

  3. Build two eyes from one image

    A shader moves each pixel sideways by its depth, half the distance for each eye. Where two pixels land on the same spot, the nearer one wins. The Screen plane slider sets which depth sits on the screen; nearer things pop out.

  4. Fill what the second eye would see

    Moving a foreground object uncovers background the camera never saw. We fill those gaps from the neighbouring background, which is fast but can smear on wide gaps.

  5. Pack for the viewer

    Red/cyan for glasses, side by side for headsets, or "wiggle": the two eyes alternate so you can see depth without any glasses.

Limits

  • Speed depends on your graphics card. Without WebGPU the page falls back to the CPU, which is much slower.
  • The 3D output updates at the depth model's rate, so it can trail the input by a frame or two.
  • Thin structures, hair and fast motion are where depth models make mistakes; text overlays should really stay flat.
  • This is a demonstration. Our server version, measured on an RTX 4070, is on the lab page.