Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints
Abstract
Building high-fidelity digital twins of articulated objects fromvisual data remains a central challenge. Existing approaches depend onmulti-view captures of the object in discrete, static states, which severelyconstrains their real-world scalability. In this paper, we introduce Articu-lat3D, a novel framework that constructs such digital twins from casuallycaptured monocular videos by jointly enforcing explicit 3D geometric andmotion constraints. We first propose Motion Prior–Driven Initialization,which leverages 3D point tracks to exploit the low-dimensional structureof articulated motion. By modeling scene dynamics with a compact set ofmotion bases, we facilitate soft decomposition of the scene into multiplerigidly-moving groups. Building on this initialization, we introduce Ge-ometric and Motion Constraints Refinement, which enforces physicallyplausible articulation through learnable kinematic primitives parameter-ized by a joint axis, a pivot point, and per-frame motion scalars, yieldingreconstructions that are both geometrically accurate and temporally co-herent. Extensive experiments demonstrate that Articulat3D achievesstate-of-the-art performance on synthetic benchmarks and real-worldcasually captured monocular videos, significantly advancing the feasibilityof digital twin creation under uncontrolled real-world conditions. Here isour project page.