Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Abstract
Online 3D reconstruction models perform poorly on longvideos. This happens because regressing poses relative to a fixed first-frameanchor forces extrapolation far beyond the training distribution. Smalldrifts accumulate and amplify into significant geometric collapse. However,we observe that per-frame depth remains stable throughout this failure.The backbone’s local geometry remains intact; only the global pose headbreaks down. Motivated by this decoupling, we introduce Scal3R. Thisapproach reformulates online reconstruction as multi-reference relativepose querying. We use lightweight learnable tokens, which make upabout ∼1% of the parameters, and inject them into a completely frozenbackbone via asymmetric attention. This setup queries poses relative tomultiple past keyframes. An online pose-graph optimization system withloop closure suppresses long-range drift. Scal3R reaches convergence in8 hours on a single GPU. It reduces the average ATE by over 60% onKITTI compared to the online baseline. It also achieves state-of-the-artperformance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and7-Scenes. Project page: https://linjohnss.github.io/scal3r/