Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
Abstract
Novel View Synthesis (NVS) enables the generation of un-seen views of a scene from a single or multiple images, allowing users tofreely explore an object from any viewpoint. Despite the recent impres-sive qualitative improvements of generative models for this task, existingmethods struggle to provide global and intuitive control of target view-points because they either use input-relative camera poses or are limitedto generating sparse global views. This lack of global pose control severelylimits the number of downstream tasks potentially enabled by NVS. Toaddress this limitation, we propose a novel approach for precise cameracontrol in a customizable Normalized Object Coordinate Space (NOCS),requiring single or few unposed images. Our method operates solely onthe absolute camera pose of the target view in NOCS, eliminating theneed for a relative world frame or camera poses of the input images. Un-like previous methods that treat NVS as a standalone generation task, weformulate it as an image editing problem and build upon state-of-the-artediting models to leverage their superior generalization capability. Cam-era information is injected as dedicated camera tokens via an in-contextmulti-modal conditioning strategy. To alleviate the inherent ambiguityof NOCS, we incorporate text descriptions that explicitly define the ob-ject’s canonical coordinate frame, which also enhances generalization tounseen object categories. Furthermore, we curate a high-quality datasetwith consistently aligned orientations and corresponding NOCS text def-initions. Extensive experiments demonstrate that our method robustlygenerates novel views with accurate and consistent orientations fromarbitrary unposed images across diverse categories, achieving state-of-the-art image quality and fidelity.