MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction
Abstract
Reconstructing articulated 3D objects from a single imagerequires jointly inferring object geometry, part structure, and motionparameters from limited visual evidence. A key difficulty lies in theentanglement between motion cues and object structure, which makesdirect articulation regression unstable. Existing methods address thischallenge through multi-view supervision, retrieval-based assembly, orauxiliary video generation, often sacrificing scalability or efficiency. Wepresent MonoArt, a unified framework grounded in progressive struc-tural reasoning. Rather than predicting articulation directly from im-age features, MonoArt progressively transforms visual observations intocanonical geometry, structured part representations, and motion-awareembeddings within a single architecture. This structured reasoning pro-cess enables stable and interpretable articulation inference without ex-ternal motion templates or multi-stage pipelines. Extensive experimentson PartNet-Mobility demonstrate that MonoArt achieves state-of-the-artperformance in both reconstruction accuracy and inference speed. Theframework further generalizes to robotic manipulation and articulatedscene reconstruction.