SynLF: Zero-Shot Metric Depth from Light Field Cameras via Physics-Grounded Synthesis
Abstract
Accurate monocular metric depth estimation remains inher-ently ill-posed due to scale ambiguity. Light field cameras mitigate thisambiguity by capturing micro-baseline disparities within a single lensto recover metric depth, but current learning-based solutions are lim-ited by scarce training data and often overfit Lambertian appearancepatterns, restricting real-world robustness. To address this, we proposeSynLF, which alleviates data scarcity via physics-grounded light field(PG-LF) synthesis and performs robust real-world metric depth esti-mation through VisDepth, a prior-initialized iterative estimator with ex-plicit multi-view occlusion reasoning. During training, we synthesize lightfield supervision on-the-fly from RGB-D via physics-grounded light fieldsynthesis, including geometry-consistent view synthesis, depth-dependentdefocus, and non-Lambertian perturbations. At inference, the model isdirectly evaluated on real light field captures without fine-tuning. Ex-tensive experiments demonstrate that SynLF generalizes zero-shot toour custom real-world light field dataset, which features highly chal-lenging conditions such as transparency, reflections, specular highlights,and textureless regions. It achieves 30 mm absolute and 1.4% relativeerrors across 0.5–2.5 m, demonstrating that physics-grounded LF syn-thesis from large-scale RGB-D provides a scalable and storage-efficientalternative to conventional light field acquisition for accurate single-lensmetric depth estimation.