MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
Abstract
Recent Vision-Language Models (VLMs) have achieved re-markable success in visual understanding, driven by the growing avail-ability of high-quality image-text pairs. However, the performance ofVLMs often degrades in the video domain due to the essential needfor temporal understanding and the scarcity of large-scale annotatedvideo data. In this work, we focus on marine video understanding, whichbrings further challenges: first, it requires substantial domain expertise;and video VLMs usually struggle with localizing and interpreting criti-cal information from marine videos, as the informative events are typ-ically sparse, unpredictable, and unevenly distributed. To address thesechallenges, we carefully curate the first event-centric marine video un-derstanding dataset called MarineEVT, which features 20K multi-task,video-level visual question-answering pairs spanning multiple dimensionsof marine understanding and analysis. Meanwhile, based on MarineEVT,we decompose marine video understanding as an Event-centric VisualTool-integrated Reasoning process (EVT-R1 for short), where we lever-age powerful visual tools to drive the model to localize and interpretcritical information aligned with visual questions and human intent. Todemonstrate its effectiveness, we compare EVT-R1 against 11 SOTAVLMs in different settings. EVT-R1 outperforms the top open-sourceand top commercial models by 5.22 and 11.09, respectively. MarineEVTand EVT-R1 lay the foundation for ecological discovery and marine ed-ucation, fostering the development of VLMs capable of interpreting ma-rine dynamics, reasoning about ecological interactions, and supportingsustainable ocean video understanding and analysis.