StreetForward: Perceiving Dynamic Street with Feedforward Causal Dynamics
Abstract
Feedforward reconstruction is crucial for autonomous drivingapplications, where rapid scene reconstruction enables efficient utiliza-tion of large-scale driving datasets in closed-loop simulation and otherdownstream tasks, eliminating the need for time-consuming per-scene op-timization. We present StreetForward, a pose-free and tracker-free feed-forward framework for dynamic street reconstruction. Building uponthe alternating attention mechanism from Visual Geometry GroundedTransformer (VGGT), we propose a simple yet effective temporal maskattention module that captures dynamic motion information from im-age sequences and produces motion-aware latent representations. Staticcontent and dynamic instances are represented uniformly with 3D Gaus-sian Splatting, and are optimized jointly by cross-frame rendering withspatio-temporal consistency, allowing the model to infer per-pixel ve-locities and produce high-fidelity novel views at new poses and times.We train and evaluate our model on the Waymo Open Dataset, demon-strating superior performance on novel view synthesis and depth estima-tion compared to existing methods. Furthermore, zero-shot inference onCARLA validates the generalization capability of our approach.