LumiDepth: Stable Monocular Depth in Multi-Illumination Scenes
Abstract
Depth estimation in multi-illumination scenes with multi-ple, spatially varying light sources remains a crucial yet less-exploredproblem. Illumination changes introduce shadows, specular highlights,and exposure shifts that distort local appearance cues, causing severedepth inconsistency or even failure. Existing depth foundation models,trained predominantly on uniformly lit data, degrade sharply under suchconditions. However, direct adaptation is challenging because groundtruth depth is typically limited for multi-illumination datasets, whilesynthetic relighting often incurs geometric distortions. To address thesechallenges, we propose LumiDepth, a framework that learns from multi-illumination RGB images. First, a Disagreement-Calibrated ProbabilisticPseudo Supervision (DCPS) module constructs high-quality pseudo la-bels while preserving diversity. Second, a Frequency-aware Consistencyand Distillation (FaCD) module improves cross-illumination stabilitywithout over-smoothing by enforcing low-frequency geometric consis-tency and distilling high-frequency structural details bi-directionally. Toenable systematic evaluation, we introduce ReMID, a real-world multi-illumination RGB-D benchmark, together with stability metrics thatquantify average and worst-case depth variation. Experiments acrossdiverse datasets demonstrate that LumiDepth achieves state-of-the-artoverall performance, markedly improving both consistency and accuracyby reducing depth variation by 30.2% and absolute relative error by24.8%. We further show our target-domain label-free design remains ef-fective for depth under other appearance shifts such as weather and sen-⋆sor noise.