Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding
Abstract
The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward3D foundation models. However, their inherent reliance on static-sceneassumptions leads to severe misalignment of spatial features in uncon-strained dynamic environments. To bridge this critical gap, we proposeSPAR, a novel joint semantic-geometric encoding architecture that ex-plicitly isolates transient dynamic noise prior to latent space aggregation.Furthermore, we introduce a dynamic-region-aware end-to-end trainingparadigm that structurally couples motion estimation with multi-viewvisual and semantic learning. This unified approach enables the networkto inherently resolve motion conflicts and distill multi-view consistent,temporally stable scene representations from dynamic inputs. Extensiveexperiments on the challenging D-RE10K benchmark demonstrate thatSPAR achieves state-of-the-art performance. Our end-to-end approachachieves exceptional novel view synthesis quality, yielding a PSNR of22.15 dB and 23.33 dB given only 3 and 4 input views respectively. De-spite being trained in a self-supervised manner, our model achieves anmIoU of 88.5% for motion mask prediction. Furthermore, our analysisreveals a strong inter-task synergy between photometric scene recon-struction and semantic understanding, where semantic synthesis learningconsistently enhances photometric fidelity in novel view rendering. Codewill be available at https://github.com/dmucby/SPAR.