Context-Interactive Reasoning for Group Activity Detection
Abstract
Group Activity Detection (GAD) aims to jointly infer groupmemberships and collective activities from videos. Existing approachestypically rely on actor-centric features to model group activities, assum-ing that collective semantics can be inferred solely from inter-actor re-lations. This overlooks critical context-interactive information, such asexplicit spatial constraints and actor–scene dependencies, which are es-sential in crowded and multi-view environments. To address these lim-itations, we propose a Context-Interactive Reasoning framework thatjointly models spatial topology and group semantic context. Specifically,a Short-Term Interaction Regularizer (STIR) softly constrains inter-actor spatial relations with learnable frame-level topology priors, sup-pressing spurious connections and promoting spatially coherent group-ing. Complementarily, a Long-Range Context Conditioner (LRCC) se-lectively incorporates global scene semantics via data-dependent gating,enabling activity-aware context utilization while avoiding unnecessarynoise. Extensive experiments on challenging benchmarks demonstratethat our method achieves state-of-the-art performance in both groupmembership identification and collective activity recognition.