GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
Abstract
Novel view synthesis requires strong 3D geometric consis-tency and the ability to generate visually coherent images across di-verse viewpoints. While recent camera-controlled video diffusion modelsshow promising results, they often suffer from geometric distortions andlimited camera controllability. To overcome these challenges, we intro-duce GeoNVS, a geometry-grounded novel-view synthesizer that en-hances both geometric fidelity and camera controllability through ex-plicit 3D geometric guidance. Our key innovation is the Gaussian Splat-ting Feature Adapter (GS-Adapter), which lifts input-view diffusion fea-tures into 3D Gaussian representations, renders geometry-constrainednovel-view features, and adaptively fuses them with diffusion features tocorrect geometrically inconsistent representations. Unlike prior methodsthat inject geometry at the input level, GS-Adapter operates in featurespace, avoiding view-dependent color noise that degrades structural con-sistency. Its plug-and-play design enables zero-shot compatibility withdiverse feed-forward geometry models without additional training, andcan be adapted to other video diffusion backbones. Experiments across 9scenes and 18 settings demonstrate state-of-the-art performance, achiev-ing 11.3% and 14.9% improvements over SEVA and CameraCtrl, withup to 2× reduction in translation error and 7× in Chamfer Distance.