首个水上环境视频问答基准,让无人船懂规则、会推理。
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
- 用多智能体系统融合视觉与规则,实现动态水域认知
- 在6类水域3029段视频上验证,显著超越现有方法
- 适合研究无人船智能决策与可解释性系统的学者
尽管自主导航在被动感知(如目标检测与分割)方面取得显著进展,但其仍受限于缺乏知识驱动的交互式环境认知。在高风险的海上航行领域,将原始视觉感知与复杂认知推理相衔接,不仅是性能提升,更是无人水面艇(ASV)执行安全精准操作的关键前提。为此,我们提出WaterVideoQA,首个专为全水路环境设计的大规模、综合性视频问答基准。该基准涵盖6种不同水路类别、共3029段视频片段,融合了变化光照与动态天气等多重变量,严格测试ASV在五级分层认知框架下的能力。此外,我们引入NaviMind,一种开创性的多智能体神经符号系统,用于开放式海上推理。通过自适应语义路由、情境感知层次推理与自主自我验证机制,NaviMind使ASV从表面模式匹配跃升至符合航行规则、可解释的决策水平。实验表明,该框架显著超越现有基线,确立了动态海事环境中智能可信交互的新范式。
原文摘要 · Abstract (English)
While autonomous navigation has achieved remarkable success in passive perception (e.g., object detection and segmentation), it remains fundamentally constrained by a void in knowledge-driven, interactive environmental cognition. In the high-stakes domain of maritime navigation, the ability to bridge the gap between raw visual perception and complex cognitive reasoning is not merely an enhancement but a critical prerequisite for Autonomous Surface Vessels to execute safe and precise maneuvers. To this end, we present WaterVideoQA, the first large-scale, comprehensive Video Question Answering benchmark specifically engineered for all-waterway environments. This benchmark encompasses 3,029 video clips across six distinct waterway categories, integrating multifaceted variables such as volatile lighting and dynamic weather to rigorously stress-test ASV capabilities across a five-tier hierarchical cognitive framework. Furthermore, we introduce NaviMind, a pioneering multi-agent neuro-symbolic system designed for open-ended maritime reasoning. By synergizing Adaptive Semantic Routing, Situation-Aware Hierarchical Reasoning, and Autonomous Self-Reflective Verification, NaviMind transitions ASVs from superficial pattern matching to regulation-compliant, interpretable decision-making. Experimental results demonstrate that our framework significantly transcends existing baselines, establishing a new paradigm for intelligent, trustworthy interaction in dynamic maritime environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。