SegFly: A 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale
Abstract
Semantic segmentation for uncrewed aerial vehicles (UAVs) isfundamental for aerial scene understanding, yet existing RGB and RGB-Tdatasets remain limited in scale, diversity, and annotation efficiencydue to the high cost of manual labeling and the difficulties of accurateRGB-T alignment on off-the-shelf UAVs. To address these challenges, wepropose a scalable geometry-driven 2D-3D-2D paradigm that leveragesmulti-view redundancy in high-overlap aerial imagery to automaticallypropagate labels from a small subset of manually annotated RGB imagesto both RGB and thermal modalities within a unified framework. Bylifting less than 3% of RGB images into a semantic 3D point cloud andrendering it into all views, our approach enables dense pseudo ground-truth generation across large image collections, automatically producing97% of RGB labels and 100% of thermal labels while achieving 91% and88% annotation accuracy without any 2D manual refinement. We furtherextend this 2D–3D–2D paradigm to cross-modal image registration, using3D geometry as an intermediate alignment space to obtain fully automatic,strong pixel-level RGB-T alignment with 87% registration accuracy and nohardware-level synchronization. Applying our framework to existing geo-referenced aerial imagery, we construct SegFly, a large-scale benchmarkwith over 20,000 high-resolution RGB images and more than 15,000geometrically aligned RGB-T pairs spanning diverse urban, industrial, andrural environments across multiple altitudes and seasons. On SegFly, weestablish the Firefly baseline for RGB and thermal semantic segmentationand show that both conventional architectures and vision foundationmodels benefit substantially from SegFly supervision, highlighting thepotential of geometry-driven 2D-3D-2D pipelines for scalable multi-modalaerial scene understanding. The SegFly dataset and our Firefly baselineare available at https://github.com/markus-42/SegFly.