Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics
Abstract
Intrinsic image decomposition aims to estimate physicallybased rendering (PBR) parameters such as albedo, roughness, and metal-licity from images. While recent methods achieve strong single-view pre-dictions, applying them independently to multiple views of the samescene often yields inconsistent estimates, limiting their use in down-stream applications such as editable neural scenes and 3D reconstruc-tion. Video-based models can improve cross-frame consistency but re-quire dense, ordered sequences and substantial compute, limiting theirapplicability to sparse, unordered image collections. We propose Geo-ID, a novel test-time framework that repurposes pretrained single-viewintrinsic predictors to produce cross-view consistent decompositions bycoupling independent per-view predictions through sparse geometric cor-respondences that form uncertainty-aware consensus targets. Geo-ID ismodel-agnostic, requires no retraining or inverse rendering, and appliesdirectly to off-the-shelf intrinsic predictors. Experiments on syntheticbenchmarks and real-world scenes demonstrate substantial improvementsin cross-view intrinsic consistency as the number of views increases, whilemaintaining comparable single-view decomposition performance. We fur-ther show that the resulting consistent intrinsics enable coherent appear-ance editing and relighting in downstream neural scene representations.