Show the numbers
| Training step | Before the fix | After the fix |
|---|---|---|
| 1,000 | 11.11 | 14.27 |
| 2,000 | 10.54 | 14.36 |
| 3,000 | 10.64 | 14.75 |
| 4,000 | 9.67 | 14.66 |
| 5,000 | 9.56 | 14.84 |
| 6,000 | 9.75 | 14.79 |
| 7,000 | 10.33 | 14.84 |
| 8,000 | 9.20 | 14.76 |
| 9,000 | 9.48 | 15.22 |
| 10,000 | 9.40 | 15.14 |
| 11,000 | 9.50 | 15.07 |
| 12,000 | 9.56 | 15.14 |
| 13,000 | 9.44 | 15.00 |
| 14,000 | 9.73 | 14.98 |
| 15,000 | 9.78 | 15.14 |
| 16,000 | 9.90 | 15.05 |
| 17,000 | 9.37 | 15.12 |
| 18,000 | 9.46 | 15.10 |
| 19,000 | 9.65 | 15.16 |
| 20,000 | 9.79 | 15.19 |
A model that turns a few phone photos of any car into a 3D splat in one pass. Our first training run got worse for 19,000 straight steps. The cause was a PyTorch function silently returning wrong gradients for the memory layout our renderer produces. After a one-line fix, the same model trains properly.
What we tried
- 31 cars from 3DRealCar (Apache-2.0): phone photos, LiDAR depth and camera poses, read straight out of a 47 GB archive without downloading it.
- A small network trained from scratch, so no non-commercial weights are inherited.
- Scored on cars it has never seen, not on unseen photos of cars it has, so it cannot pass by memorising.
What we measured
| Measure | Before the fix | After the fix | Note |
|---|---|---|---|
| Best score on 5 unseen cars | 11.1 dB | 15.22 dB | AnySplat: 14.6 |
| Score at the end of training | 9.8 dB | 15.19 dB | |
| One fixed image pair, 600 steps | 14.7 → 9.5 dB | 14.7 → 30.2 dB | should always improve |
| Gradient check on the SSIM term | 0.026 | 1.001 | 1.00 means the gradient is right |
| Training time, one RTX 4070 | — | 30 min |
What went wrong
- The renderer returns images as height × width × colour. Reordering that to colour × height × width creates a 'view' with an unusual memory layout, not a copy.
- On our PyTorch version (2.4.1, CUDA), the backward pass of avg_pool2d returns wrong gradients for that layout: up to 0.70 off per pixel, with no error or warning. Our image-similarity loss (SSIM) used it.
- So every training step followed a corrupted signal. The model learned to hide the damage with larger, fainter Gaussians, which is why its renders grew blurrier.
- Finding it took elimination: cameras, depth, rotation maths, the optimiser and the network were each tested and cleared before the loss term itself was. The fix is one line, .contiguous() before pooling.
- Our reflection-capture code was checked for the same fault. Its SSIM uses a different operation and passes the gradient check at 0.999.
What happens next
- Let the input photos exchange information before placing points, which is what the strongest 360° methods do.
- Re-test SSIM now that its gradient is fixed, and train on more cars.
Built with
- 3DRealCar dataset Apache-2.0
- gsplat 1.5.3 Apache-2.0
- PyTorch 2.4.1 BSD-3
Next experiment
A housing society in 3D, with every plot's owner, file and payments on top →
Want this for your data?