RGB-Pointmap Pretraining for Unified 3D Scene Understanding
Abstract
Pretraining 3D encoders through alignment with ContrastiveLanguage–Image Pre-training (CLIP) has emerged as a promising direc-tion to learn generalizable representations for 3D scene understanding. Inthis paper, we propose UniScene3D, a transformer-based framework thatlearns unified 3D scene representations from multi-view RGB–Pointmapby leveraging the priors of a pretrained 2D foundation model. For robustRGB-Pointmap representation learning, we introduce novel cross-viewgeometric alignment and grounded view alignment to enforce geome-try and semantic consistency across views. Extensive low-shot and task-specific fine-tuning across viewpoint grounding, scene retrieval, sceneclassification, and 3D visual question answering achieves state-of-the-artperformance. These results establish UniScene3D as an effective frame-work for unified 3D scene understanding.