Video Generative Models as Geometry Learner
Abstract
Recent generative approaches to geometry estimation adaptpretrained image diffusion models and treat the task as image-conditionedgeneration. Leveraging off-the-shelf image diffusion models, they either(i) train task-specific geometry models (for depth and surface normal es-timation) independently, losing the opportunity of exploring the intrinsiccorrelation of these geometric targets, or (ii) jointly fine-tune modifiedimage diffusion backbones (e.g., altered self-attention), which typicallydemands substantial labeled data. To overcome these limitations in aprincipled fashion, we repurpose pretrained video generative models as aunified and data-efficient framework for geometry estimation, formulatedinnovatively as a next-frames prediction task. Our method, GeoNeXt,inherits naturally structured knowledge and richer priors from the videomodel, while further adapting them for joint modeling of images and ge-ometry targets (