Do Multimodal LLMs Understand Intraoral Dental Data? Dataset, Platform, and Baselines
Abstract
Progress in dental computer vision is limited by the absence of large-scale multimodal datasets that jointly capture 3D intraoral geometry and 2Dappearance across diverse clinical settings. Existing resources are typically uni-modal, which hinders robust cross-modal learning and generalization. We as-semble and release a multi-center dataset of 1,000 patients comprising 2,000registered upper/lower intraoral scans, 5,000 paired intraoral photographs, and2,403 clinician-authored reports. This combination links detailed 3D dental ge-ometry with complementary 2D evidence, supporting occlusal and orthodonticanalysis. Moreover, to enable scalable and privacy-preserving acquisition andannotation across distributed centers, we introduce an open platform that sup-ports multimodal ingestion and structured labeling. Experiments indicate thatstate-of-the-art multimodal models fail to generate clinically faithful reports, moti-vating geometry-aware adaptation. We therefore propose IOS-Qwen, which fuses aPointTransformer 3D encoder with Qwen3-VL to generate structured, point-cloud-conditioned reports. Together, the dataset, the platform, and the baselines establisha foundation for multimodal dental AI research. Code is publicly released.3