← Lab bench
EXP-049Result

The silent PyTorch bug that made our car model get worse

15.22dB

on 5 cars it has never seen, above the released AnySplat model (14.6), trained from scratch on commercially licensed data

Score on 5 cars the model never saw, during trainingHover or use arrow keys
81012141605k10k15k20kTraining stepPSNR, dBReleased model (AnySplat) 14.6Untrained, depth sensor only 12.0Before the fixAfter the fix: best 15.22 dB
Show the numbers
Training stepBefore the fixAfter the fix
1,00011.1114.27
2,00010.5414.36
3,00010.6414.75
4,0009.6714.66
5,0009.5614.84
6,0009.7514.79
7,00010.3314.84
8,0009.2014.76
9,0009.4815.22
10,0009.4015.14
11,0009.5015.07
12,0009.5615.14
13,0009.4415.00
14,0009.7314.98
15,0009.7815.14
16,0009.9015.05
17,0009.3715.12
18,0009.4615.10
19,0009.6515.16
20,0009.7915.19

A model that turns a few phone photos of any car into a 3D splat in one pass. Our first training run got worse for 19,000 straight steps. The cause was a PyTorch function silently returning wrong gradients for the memory layout our renderer produces. After a one-line fix, the same model trains properly.

What we tried

  • 31 cars from 3DRealCar (Apache-2.0): phone photos, LiDAR depth and camera poses, read straight out of a 47 GB archive without downloading it.
  • A small network trained from scratch, so no non-commercial weights are inherited.
  • Scored on cars it has never seen, not on unseen photos of cars it has, so it cannot pass by memorising.

What we measured

MeasureBefore the fixAfter the fixNote
Best score on 5 unseen cars11.1 dB15.22 dBAnySplat: 14.6
Score at the end of training9.8 dB15.19 dB
One fixed image pair, 600 steps14.7 → 9.5 dB14.7 → 30.2 dBshould always improve
Gradient check on the SSIM term0.0261.0011.00 means the gradient is right
Training time, one RTX 4070—30 min

What went wrong

  • The renderer returns images as height × width × colour. Reordering that to colour × height × width creates a 'view' with an unusual memory layout, not a copy.
  • On our PyTorch version (2.4.1, CUDA), the backward pass of avg_pool2d returns wrong gradients for that layout: up to 0.70 off per pixel, with no error or warning. Our image-similarity loss (SSIM) used it.
  • So every training step followed a corrupted signal. The model learned to hide the damage with larger, fainter Gaussians, which is why its renders grew blurrier.
  • Finding it took elimination: cameras, depth, rotation maths, the optimiser and the network were each tested and cleared before the loss term itself was. The fix is one line, .contiguous() before pooling.
  • Our reflection-capture code was checked for the same fault. Its SSIM uses a different operation and passes the gradient check at 0.999.

What happens next

  • Let the input photos exchange information before placing points, which is what the strongest 360° methods do.
  • Re-test SSIM now that its gradient is fixed, and train on more cars.

Built with

  • 3DRealCar dataset Apache-2.0
  • gsplat 1.5.3 Apache-2.0
  • PyTorch 2.4.1 BSD-3