AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting
Abstract
While 3D Vision Foundation Models (3DVFMs) have demon-strated remarkable zero-shot capabilities in visual geometry estimation,their direct application to generalizable novel view synthesis (NVS) re-mains challenging. In this paper, we propose AirSplat, a novel train-ing framework that effectively adapts the robust geometric priors of3DVFMs into high-fidelity, pose-free NVS. Our approach introduces twokey technical contributions: (1) Self-Consistent Pose Alignment (SCPA),a training-time feedback loop that ensures pixel-aligned supervision toresolve pose-geometry discrepancy; and (2) Rating-based Opacity Match-ing (ROM), which leverages the local 3D geometry consistency knowl-edge from a sparse-view NVS teacher model to filter out degraded prim-itives. Experimental results on large-scale benchmarks demonstrate thatour method significantly outperforms state-of-the-art pose-free NVS ap-proaches in reconstruction quality. Our AirSplat highlights the potentialof adapting 3DVFMs to enable simultaneous visual geometry estimationand high-quality view synthesis.