RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception
Abstract
LiDAR has increasingly been integrated into traffic cam-eras to expand coverage and mitigate occlusion in roadside cooperativeperception. However, how unimodal and camera–LiDAR fusion archi-tectures behave under variations in LiDAR point sparsity induced bysensor configurations and scene-dependent sensing conditions remainsunderexplored. We introduce RESOLVE, a large-scale real-world bench-mark dataset featuring multi-resolution roadside LiDAR and synchro-nized camera-LiDAR sensing for systematic evaluation of unimodal andfusion-based architectures in roadside 3D detection and tracking. RE-SOLVE contains over 100k images and 26k point cloud frames with 220kmanually annotated bounding boxes, captured at a real-world urban in-tersection across diverse lighting and weather conditions and spanning10 classes of traffic participants. In particular, RESOLVE enables con-trolled evaluation across three LiDAR resolution levels while keeping allother sensing and environmental factors fixed. This allows fair cross-architecture comparisons under point cloud distribution shifts resultingfrom resolution variations, sensing distance, and training–inference res-olution mismatches. Results from extensive benchmark experiments re-veal insights into how multimodal fusion can compensate for LiDARpoint sparsity, offering clues for designing cost-efficient roadside multi-modal perception. The dataset and benchmark codes are available athttps://github.com/ASU-Suo-Lab/RESOLVE.