Diversity-Aware View Partitioning for Scalable VGGT
Abstract
Geometry transformers such as VGGT achieve strong per-formance by jointly reasoning over multiple views with global attention.However, scaling them to large view collections remains challenging dueto the quadratic cost of attention. Moreover, our empirical analysis re-veals that the reconstruction quality in VGGT is sensitive to the distri-bution of viewpoints. Simply increasing the number of views without suf-ficient viewpoint diversity can even degrade performance, as redundantviews introduce highly similar tokens that dilute informative geometricsignals in the attention mechanism. Motivated by this observation, wepropose a training-free and plug-and-play VGGT inference frameworkthat organizes views into diversity-aware balanced chunks. The chunksare constructed through combinatorial graph partitioning over visualdissimilarity and spatial dispersion. This view organization allows thetransformer to focus attention on geometrically informative views whilereducing redundant attention interactions. To estimate spatial dispersionwithout full pose estimation, we approximate spatial relationships via asoft pose propagation strategy based on visual similarity from a smallset of seed frames. Extensive experiments demonstrate improved perfor-mance in camera pose estimation, multi-view depth prediction, and 3Dreconstruction while reducing memory usage and inference latency. Ourframework also complements existing VGGT variants, enabling scalablemulti-view reconstruction without sacrificing geometric fidelity.