LVSPM: Long Sequence View Synthesis and Pose Estimation Model
Abstract
We present LVSPM, a generalizable model that jointly esti-mates camera poses and synthesizes novel views from uncalibrated im-age collections. Trained with only RGB images and pose supervision,LVSPM avoids dense 3D ground truth and employs test-time train-ing (TTT) layers to scale seamlessly to hundreds of input views. OnRealEstate10k, Co3Dv2, and DL3DV, LVSPM surpasses VGGT in poseestimation across 16–256 views, with especially large margins at strictthresholds. For novel view synthesis under a practical protocol wheremore views cover larger scenes, LVSPM achieves state-of-the-art pose-free quality—surpassing even pose-dependent models in PSNR—and stillmaintains high quality as scene scale grows, while baselines collapse. Thecode will be available at https://burningdust21.github.io/Projects/LVSPM/.