MCVL: Multi-Space Cross-View Learning for Aerial-Ground Person Re-Identification
Abstract
The core challenge of Aerial-Ground Person Re-Identification (AG-ReID) is the drastic viewpoint discrepancy, which induces a significant cross-view feature gap that hinders aerial-ground matching. However, existing AG-ReID methods overlook this gap and attempt to learn direct global semantic correspondences through auxiliary components (e.g., attributes, tokens, or prompts), which limits training effectiveness and increases computational overhead. To address these issues, we propose a novel AG-ReID framework, MCVL, which effectively bridges the cross-view gap through multi-space cross-view learning. Our method aligns both spatial and global feature spaces, learning view-invariant yet identity-discriminative representations to effectively handle viewpoint discrepancies. Specifically, we introduce a Deformable Homography Transformation (DHT) module that projects features across aerial and ground view planes, creating an intermediate cross-view spatial feature space. Based on this, we propose a Cross-View Learning (CVL) strategy that leverages both original and projected features to reduce the spatial feature gap and enhance identity discrimination while preserving perspective invariance. Additionally, we design a View Decorrelation Contrastive Loss (VDCL) to mitigate view-specific biases by pushing global features away from their corresponding view-specific prototypes while pulling together global features of the same identity, thereby suppressing view-related cues and strengthening identity-related representations. Extensive experiments on five AG-ReID datasets, including CARGO, AG-ReIDv1, AG-ReIDv2, LAGPeR, and G2APS-ReID, demonstrate the effectiveness of our proposed MCVL method.