Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
Abstract
Recent 3D generative models produce high-quality geometryfrom a single image using large-scale priors and diffusion architectures.However, they assume complete visibility and single-view inputs, limitingapplicability in multi-view, occluded, or editing scenarios. Although priorworks address these challenges individually, they lack a unified frameworkfor controllable 3D completion under diverse conditioning signals.We present Axolotl3D, a multi-modal and occlusion-aware 3D generationmodel that jointly conditions on images, visibility masks, camera param-eters, and a partial point cloud. The point cloud serves as a geometricanchor promoting faithful shape completion, while camera parametersensure consistent multi-view alignment in a shared 3D coordinate sys-tem. A unified training strategy synthesizes diverse conditioning regimesfrom large-scale 3D data, enabling robust cross-modal reasoning.Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strongresults in real-world reconstruction and geometry-consistent editing.