arXiv:2608.08281cs.AIcs.RO2026-08中稿 · the IFAC WC 2026 a…

用大模型理解海上航行规则与实际操作,验证其在真实场景中的推理能力。

Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios

论文配图:Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios
图 1 · 摘自论文原文
  • 构建50个真实航行动态数据集,标注碰撞规则与最佳船艺建议。
  • 大模型在未微调情况下仍难以准确判断航行规则和应对策略。
  • 适合研究智能航海、多模态决策的学者与开发者参考。

近期大型语言模型(LLMs)在多个领域展现出强大的情境理解、推理与决策能力,尤其在汽车领域表现突出。本文探索当前最先进的LLMs在海上航行任务中的应用,涵盖《碰撞规则》(COLREGs)中的明文规定以及“良好船艺”(Good Seamanship)所概括的非成文实践。我们基于AIS数据构建了包含50个多样化真实航行场景的数据集,对每个场景标注适用的COLREG规则、推荐行动及行动依据。通过测试多种不同架构与规模的LLM,评估其在航海任务中的理解能力与推理表现。结果表明,即使在较大在线模型中,未经微调也难以有效完成航海任务,说明该领域仍具挑战性。

原文摘要 · Abstract (English)

Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations (COLREGs) and uncodified best practices summarized in the concept of ``Good Seamanship''. We construct a dataset consisting of 50 diverse, real-world navigation scenarios from AIS data, label scenarios with applicable COLREG rules, recommended actions, and the reasoning for the action. We explore a variety of different LLM architectures and sizes to determine their understanding of maritime navigation tasks as well as evaluate their reasoning capabilities in this domain. The results obtained indicate that the maritime navigation task remains difficult to solve without fine-tuning, even for larger online models.

大模型航海安全推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。