Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
Abstract
Multimodal Large Language Models (MLLMs) have demon-strated immense potential in Earth observation. However, the massivevisual tokens generated when processing Ultra-High-Resolution (UHR) im-agery introduce prohibitive computational overhead, severely bottleneck-ing their inference efficiency. Existing visual token compression methodspredominantly adopt static and uniform compression strategies, neglectingthe inherent Semantic–Geometric Duality in remote sensing interpretationtasks. Specifically, object semantic tasks focus on the abstract seman-tics of objects and benefit from aggressive background pruning, whereasscene geometric tasks critically rely on the integrity of spatial topol-ogy. To address this challenge, we propose DualComp, a task-adaptivedual-stream token compression framework. Dynamically guided by alightweight pre-trained router, DualComp decouples feature processinginto two dedicated pathways. In the object semantic stream, the Spatially-Contiguous Semantic Aggregator (SCSA) utilizes size-adaptive clusteringto aggregates redundant background while protecting small object. Inthe scene geometric stream, the Instruction-Guided Structure Recoverer(IGSR) introduces a greedy path-tracing topology completion mechanismto reconstruct spatial skeletons. Experiments on the UHR remote sens-ing benchmark XLRS-Bench demonstrate that DualComp accomplisheshigh-fidelity remote sensing interpretation at an exceptionally low com-putational cost, achieving simultaneous improvements in both efficiencyand accuracy.