← Lab bench
EXP-052Live

Flat video turned into 3D live, on one graphics card

46.6fps

720p frames converted to side-by-side 3D per second on one RTX 4070. A live feed needs 30.

Robot cell: the flat video frame and the depth the model estimated for it (AI depth (light = near))Robot cell: the flat video frame and the depth the model estimated for it (Flat input)
◂▸
Flat inputAI depth (light = near)

An AI depth model estimates how far away every pixel is, and a GPU shader moves pixels sideways by their depth to build a separate left and right eye. The whole chain runs live, from video decode to hardware encode, and adds about 19 ms at 720p. We scored the AI-made right eye against a real second camera.

What we tried

  • Rendered our own test clip in Blender: three shots with hard cuts, a true right eye 65 mm away and true depth for every frame.
  • Depth Anything V2 Small per frame, with the near/far range smoothed over time and reset at every cut.
  • A GPU forward warp with a z-buffer, half the disparity per eye, and background-side hole filling.
  • FFmpeg with NVDEC in and NVENC out on three threads, so frames stay on the graphics card.
  • Video Depth Anything Small in streaming mode, as the flicker-free alternative.
  • The same model in the browser through Transformers.js on WebGPU, with the warp as a shader.

What we measured

MeasureFlat videoOur live pipelineNote
Depth error vs a real second camera10.4 px2.1 pxMean disparity error over 270 frames; AI scale set once per shot
Right-eye image match (PSNR)12.3 dB17.0 dBThe same warp fed true depth reaches 25.9 dB, so depth edges are the gap
720p throughput—46.6 fpsMedian of 3 runs (44.4–48.7); 19.9 ms GPU per frame
1080p throughput—33.5 fpsMedian of 3 runs; little headroom
Added latency, 720p live—18.8 ms median, 21.9 ms p95
GPU memory—0.29 GB at 720p
Scene cuts caught—2 of 2, no false alarms
In the browser (WebGPU)—About 37 ms depth per frameSame RTX 4070, Chromium

What went wrong

  • Feeding PNG frames capped the converter at 6.5 fps because decoding PNGs on the CPU was the bottleneck. A real H.264 feed through the hardware decoder removed it.
  • Running the depth network eagerly from Python was launch-bound. Recording it once as a CUDA graph made it about three times faster.
  • On a shared desktop GPU, settings that were only just real time fell behind whenever another app took GPU time, and latency climbed to 150–330 ms.
  • Fed over SRT, frames arrived in bursts and queued: 102 ms median latency in that test, not yet tuned.
  • Video Depth Anything flickers 14% less, but its reference streaming code ran at 15.5 fps, a third of the speed.

What happens next

  • A TensorRT engine for the depth model, aiming at two 720p feeds per card.
  • Real footage: an Insta360 clip and broadcast-style sport.
  • Overlays kept flat, better hole filling, and WebRTC delivery measured glass to glass.
  • A Unity player for Quest with a comfort test over 20 minutes.

Built with

  • Depth Anything V2 Small Apache-2.0
  • Video Depth Anything Small Apache-2.0
  • PyTorch, Hugging Face Transformers BSD-3, Apache-2.0
  • FFmpeg with NVDEC/NVENC LGPL/GPL
  • Transformers.js, ONNX Runtime Web Apache-2.0, MIT
  • Blender (test footage only) GPL