

◂▸
We walked a room for under three minutes wearing a Meta Quest 3, recording its passthrough cameras, depth sensor and head tracking with an open-source capture app. Because the headset already knows where it is, we skipped the usual camera solve: depth fusion refined the headset's own poses, and a Gaussian splat trained locally in 16 minutes. The room comes back in real metres with the floor at zero, so it can be stood in again at true scale in VR.
What we tried
- Recorded the room with OpenQuestCapture on a Quest 3: both passthrough cameras, depth and head and controller poses, 167 seconds.
- Pulled the session over Wi-Fi adb and converted it with quest-3d-reconstruction: depth fusion, pose refinement and a COLMAP export that keeps the headset's poses.
- Trained a Gaussian splat in LichtFeld Studio with exposure correction and mip anti-aliasing, capped at one million Gaussians.
- Built a web replay that flies the camera along the real head path, beside the headset's own camera clip, with an Enter VR mode at true scale.
What we measured
| Measure | Result | Note |
|---|---|---|
| Capture | 167 s, 492 frames per eye | about 2.9 frames per second, 1280 × 1280 |
| Head path walked | 38.2 m | |
| Conversion (CPU) | 9.6 min | YUV to RGB, depth fusion, pose refinement, COLMAP export |
| Training (RTX 4070) | 16 min | 30,000 steps |
| Held-out quality | 26.9 dB PSNR, SSIM 0.85 | 124 frames never used in training (every 8th) |
| Gaussians | 656,789 | |
| Download | 10.9 MB | SOG, from a 163 MB PLY |
| Head track vs camera path | 7 cm median | independent check that the poses agree |
What went wrong
- The first export kept 5 of 270 frames. A blur filter dropped 60% of the 3 fps frames, and the 'optimised' colour set only covers every 20th frame. We turned the filter off and kept the headset's poses for every frame.
- The fused depth cloud came out 6 m tall in an ordinary room: the two wall mirrors show 'rooms' behind them. We now take the floor from the headset's tracking (y = 0) instead of the points.
- Seeking the replay video failed on a simple static server without range requests, so the page now loads the clip into memory first.
What happens next
- Capture at 10 frames a second with QuestRealityCapture and compare.
- Try learned stereo depth (FoundationStereo) for cleaner geometry.
- Walk the same path in Meta's Hyperscape and our splat, and report what we can measure.
- Capture a car, a garden and an outdoor walk to find the limits of 3 m depth.
Built with
- OpenQuestCapture (capture app) MIT
- quest-3d-reconstruction (Open3D) MIT
- LichtFeld Studio (training) GPL-3.0
- Spark 2.2 + three.js r180 (viewer) MIT
Next experiment
Reading rating plates, and showing where →
Want this for your data?