arXiv:2607.24064cs.CVcs.AI2026-07中稿 · ECCV

首个面向海洋视频的事件中心理解数据集,提升模型对稀疏关键事件的定位与推理能力。

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

论文配图:MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
图 1 · 摘自论文原文
  • 构建事件中心的视觉工具推理框架,引导模型聚焦海洋视频中的关键信息
  • 在2万条多任务视频问答上,性能超越顶尖开源与商业模型5.22%和11.09%
  • 适合海洋生态研究、智能视频分析及具身认知系统开发者使用

近期视觉语言模型(VLMs)在图像理解任务中取得显著进展,得益于高质量图像-文本对的积累。然而,由于视频任务需要时序理解且大规模标注视频数据稀缺,现有VLMs在视频领域表现下降。本文聚焦海洋视频理解,面临双重挑战:需深厚领域知识,且关键事件通常稀疏、不可预测且分布不均。为此,我们精心构建首个事件中心的海洋视频理解数据集MarineEVT,包含20,000个跨维度的视频级视觉问答对。基于此,我们提出事件中心视觉工具集成推理框架EVT-R1,利用强大视觉工具辅助模型定位并解释与问题及人类意图一致的关键信息。在多种设置下对比11个SOTA VLMs,EVT-R1分别超越最佳开源与商业模型5.22%与11.09%。MarineEVT与EVT-R1为生态发现与海洋教育奠定基础,推动具备海洋动态解析、生态交互推理能力的VLM发展,支持可持续海洋视频分析。

原文摘要 · Abstract (English)

Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essential need for temporal understanding and the scarcity of large-scale annotated video data. In this work, we focus on marine video understanding, which brings further challenges: first, it requires substantial domain expertise; and video VLMs usually struggle with localizing and interpreting critical information from marine videos, as the informative events are typically sparse, unpredictable, and unevenly distributed. To address these challenges, we carefully curate the first event-centric marine video understanding dataset called MarineEVT, which features 20K multi-task, video-level visual question-answering pairs spanning multiple dimensions of marine understanding and analysis. Meanwhile, based on MarineEVT, we decompose marine video understanding as an Event-centric Visual Tool-integrated Reasoning process EVT-R1 for short, where we leverage powerful visual tools to drive the model to localize and interpret critical information aligned with visual questions and human intent. To demonstrate its effectiveness, we compare EVT-R1 against 11 SOTA VLMs in different settings. EVT-R1 outperforms the top open-source and top commercial models by 5.22 and 11.09, respectively. MarineEVT and EVT-R1 lay the foundation for ecological discovery and marine education, fostering the development of VLMs capable of interpreting marine dynamics, reasoning about ecological interactions, and supporting sustainable ocean video understanding and analysis.

视频理解海洋生态视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。