Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
Abstract
Text-driven inversion of generative models is a core paradigmfor manipulating 2D or 3D content, unlocking numerous applicationssuch as text-based editing, style transfer, or inverse problems. How-ever, it relies on the assumption that generative models remain sen-sitive to natural language prompts. We demonstrate that for state-of-the-art native text-to-3D generative models, this assumption often col-lapses. We identify a critical failure mode where generation trajectoriesare drawn into latent “sink traps”: regions where the model becomesinsensitive to prompt modifications. In these regimes, changes to theinput text fail to alter internal representations in a way that altersthe output geometry. Crucially, we observe that this is not a limita-tion of the model’s geometric expressivity; the same generative mod-els possess the ability to produce a vast diversity of shapes but, as wedemonstrate, become insensitive to out-of-distribution text guidance. Weinvestigate this behavior by analyzing the sampling trajectories of thegenerative model, and find that complex geometries can still be repre-sented and produced by leveraging the model’s unconditional generativeprior. This leads to a more robust framework for text-based 3D shapeediting that bypasses latent sinks by decoupling a model’s geometricrepresentation power from its linguistic sensitivity. Our approach ad-dresses the limitations of current 3D pipelines and enables high-fidelitysemantic manipulation of out-of-distribution 3D shapes. Project web-page: https://daidedou.sorpi.fr/publication/beyondprompts