MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
ZIREN GONG ⋅ Xiaohan Li ⋅ Fabio Tosi ⋅ Ninghui Xu ⋅ Stefano Mattoccia ⋅ Jianfei Cai ⋅ Matteo Poggi
Abstract
This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feedforward model from the 3R family to process RGB videos and regress local point maps, and on a merging model, MAGMA, that combines local maps at both intra-agent and inter-agent levels to obtain the final, global point map. Furthermore, MAGiSt3R performs pose graph optimization to mitigate cumulative camera drift occurring along the feedforward pipeline. We evaluate MAGiSt3R on both synthetic and realworld datasets, demonstrating its superior reconstruction and camera tracking accuracy compared to state-of-the-art feed-forward approaches.
Successful Page Load