arXiv:2512.24470cs.ROcs.AI2025-12

用视觉语言模型实现海上自主航行的语义避险与可控接管

Foundation models on the bridge: Semantic hazard detection and safety maneuvers for maritime autonomy with vision-language models

  • 用视觉语言模型理解海洋场景语义,识别如潜水员、火灾等异常
  • 在10秒内完成从警报到接管动作的全过程,保障人机协同安全
  • 仅靠摄像头即可运行,适合实际海事操作场景

国际海事组织(IMO)草案MASS代码要求自主或远程控制的船舶在偏离设计域时,必须进入预设应急状态、通知操作员、允许即时人工接管,并未经批准不得更改航程。在警报到接管的空窗期,需要短时程且可被人工干预的应急措施。传统海上自主系统在依赖语义判断的任务中表现不佳(如‘潜水员下水’标志意味着水中有人,‘附近起火’代表危险)。本文提出:(i) 视觉语言模型(VLMs)能提供此类分布外情况的语义感知;(ii) 采用快-慢异常检测流水线,结合短时程、人类可控的应急动作,可在交接窗口内实用化。我们提出Semantic Lookout,一种仅使用摄像头、候选动作受约束的VLM应急动作选择器,从水域有效、世界锚定的轨迹中选出一个谨慎动作(或保持位置),全程受人类持续控制。在40个港口场景中测试了每轮场景理解能力与延迟、与人类共识的一致性(模型三票多数投票)、火灾场景下的短时程风险缓解效果,以及从警报→应急动作→操作员接管的端到端流程。子10秒模型保留了大部分先进模型的语义理解能力。相比仅依赖几何信息的基线,该方法显著提升了对火灾场景的避让距离。实地测试验证了全流程可行性。结果表明,基于VLM的语义应急动作选择器符合草案MASS标准,在合理延迟范围内可行,未来可结合领域自适应与多传感器鸟瞰感知、短时重规划,发展混合自主系统。

原文摘要 · Abstract (English)

The draft IMO MASS Code requires autonomous and remotely supervised maritime vessels to detect departures from their operational design domain, enter a predefined fallback that notifies the operator, permit immediate human override, and avoid changing the voyage plan without approval. Meeting these obligations in the alert-to-takeover gap calls for a short-horizon, human-overridable fallback maneuver. Classical maritime autonomy stacks struggle when the correct action depends on meaning (e.g., diver-down flag means people in the water, fire close by means hazard). We argue (i) that vision-language models (VLMs) provide semantic awareness for such out-of-distribution situations, and (ii) that a fast-slow anomaly pipeline with a short-horizon, human-overridable fallback maneuver makes this practical in the handover window. We introduce Semantic Lookout, a camera-only, candidate-constrained VLM fallback maneuver selector that selects one cautious action (or station-keeping) from water-valid, world-anchored trajectories under continuous human authority. On 40 harbor scenes we measure per-call scene understanding and latency, alignment with human consensus (model majority-of-three voting), short-horizon risk-relief on fire hazard scenes, and an on-water alert->fallback maneuver->operator handover. Sub-10 s models retain most of the awareness of slower state-of-the-art models. The fallback maneuver selector outperforms geometry-only baselines and increases standoff distance on fire scenes. A field run verifies end-to-end operation. These results support VLMs as semantic fallback maneuver selectors compatible with the draft IMO MASS Code, within practical latency budgets, and motivate future work on domain-adapted, hybrid autonomy that pairs foundation-model semantics with multi-sensor bird's-eye-view perception and short-horizon replanning. Website: kimachristensen.github.io/bridge_policy

海上自主视觉语言模型语义感知应急接管

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。