MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
Abstract
Generative novel view synthesis faces a fundamental dilemma:geometric priors provide spatial alignment but become sparse and inaccu-rate under view changes, while appearance priors offer visual fidelity butlack geometric correspondence. Existing methods either propagate geo-metric errors throughout generation or suffer from signal conflicts whenfusing both statically. We introduce MoCam, which employs structureddenoising dynamics to orchestrate a coordinated progression from geom-etry to appearance within the diffusion process. MoCam first leveragesgeometric priors in early stages to anchor coarse structures and toleratetheir incompleteness, then switches to appearance priors in later stagesto actively correct geometric errors and refine details. This design natu-rally unifies static and dynamic view synthesis by temporally decouplinggeometric alignment and appearance refinement within the diffusion pro-cess. Experiments demonstrate that MoCam significantly outperformsprior methods, particularly when point clouds contain severe holes ordistortions, achieving robust geometry-appearance disentanglement.