OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
Abstract
Reconstructing 3D geometry from streaming video requirescontinuous inference under bounded resources. Recent geometric foun-dation models achieve impressive reconstruction quality through all-to-all attention, yet their quadratic cost confines them to short, offline se-quences. Causal-attention variants such as StreamVGGT enable single-pass streaming but accumulate an ever-growing KV cache, exhaustingGPU memory within hundreds of frames and precluding the long-horizondeployment that motivates streaming inference in the first place. Wepresent OVGGT, a training-free framework that bounds both memoryand compute to a fixed budget regardless of sequence length. Our ap-proach combines Self-Selective Caching, which leverages FFN residualmagnitudes to compress the KV cache while remaining fully compatiblewith FlashAttention, with Dynamic Anchor Protection, which shieldscoordinate-critical tokens from eviction to suppress geometric drift overextended trajectories. Extensive experiments on indoor, outdoor, andultra-long datasets show that OVGGT processes arbitrarily long videoswithin a constant VRAM envelope while achieving state-of-the-art 3D ge-ometric accuracy. The code is available at github.com/VAISR/OVGGT.