STVFocus: Query-guided Spatio-Temporal Visual Focusing for Video LLMs
Abstract
Video Large Language Models (Video LLMs) aim to under-stand and reason over dynamic visual content. However, videos inher-ently exhibit low task-relevant information density because only a smallfraction of frames and regions provide meaningful cues. As a result, theuseful signals are sparse across space and time. We present STVFocus, atraining-free framework that unifies spatio-temporal focusing to capturesalient contexts and build informative video representations. STVFocusconsists of three components. First, Hierarchical Temporal Focus appliesa coarse-to-fine frame selection strategy that captures global temporalstructure while selectively emphasizing informative moments without ex-cessive query bias. Second, Query-informed Spatial Focus uses spatialsaliency to highlight query-relevant intra-frame regions. Finally, Tem-poral Reallocation reinvests the saved spatial capacity into additionalhigh-value frames, enabling richer temporal coverage. By jointly leverag-ing temporal and spatial saliency, STVFocus produces compact yet se-mantically enriched video representations. Experiments on VideoMME,LongVideoBench, and MLVU show consistent gains with meaningfulmargins, confirming the effectiveness of unified spatio-temporal focus-ing.