arXiv:2608.19661cs.RO2026-08中稿 · the IEEE IROS 2026…

用世界模型让大模型规划航行指令更靠谱,避免撞上风电设施。

World-Model-Grounded LLM Planning for AUV and ASV Navigation Near Offshore Wind Farms

论文配图:World-Model-Grounded LLM Planning for AUV and ASV Navigation Near Offshore Wind Farms
图 1 · 摘自论文原文
  • 大模型结合物理世界模型,判断动作时长和可行性
  • 水上水下车辆均零碰撞,仿真误差降低70%以上
  • 可用卫星图自动识别障碍物,替代传感器部署

大语言模型可将自然语言任务转化为机器人动作序列,但缺乏物理感知能力:无法判断指令持续时间或是否会撞上障碍。本文提出使用世界模型增强基于大语言模型的规划器。方法包含三部分:物理驱动的神经世界模型、三阶段梯度轨迹优化器,以及带信任域保护的模型预测控制式闭环重规划器。语言模型决定做什么,世界模型决定怎么做及持续多久——无论是6自由度的8推进器AUV,还是3自由度差速驱动的ASV。在两类近海风电设施环境中的五项基准任务中,两种车辆均实现零预测碰撞。在洋流、波浪与推进器动力学下的GazeboSim仿真中,目标距离误差相比无地基基线降低70%-82%(ASV)和约93%(AUV),经残差微调后,虚拟推演的均方根误差(RMSE)分别降低60%(AUV)和69%(ASV)。对ASV进一步展示基于视觉语言模型(VLM)的语义映射流程,从卫星图像、海图和预报API中提取障碍物与环境信息,无需机载传感器,导航准确率达96%,可直接替代人工设定障碍几何。

原文摘要 · Abstract (English)

Large language models can turn a natural-language mission into a sequence of robot actions, but they do not have a sense of physics: they cannot judge how long a command should run, or whether it will make the robot drift into an obstacle. We proposed the use of a world model to expand the capabilities of Large Language model-based planners. Our method has three components: a physics-grounded neural world model, a three-phase gradient-based trajectory optimizer, and a Model Predictive Controller (MPC)-style closed-loop replanner with a trust-region guard. The language model decides what to do, and the world model decides how long, whether that means driving eight thrusters through 6 DOF or two differential thrusters through 3 DOF. We evaluate two marine vehicle classes operating near offshore wind infrastructure: a 6-DOF Autonomous Underwater Vehicle (AUV) and a 3-DOF differential-drive Autonomous Surface Vehicle (ASV). In five benchmark missions per platform, both vehicles reach every goal with zero predicted collisions, and both transfer to GazeboSim under ocean current, waves, and thruster dynamics, remaining collision-free and cutting GazeboSim goal-distance error versus the ungrounded baseline by 70-82% (ASV) and roughly 93% (AUV), after a residual fine-tuning pass that separately reduces surrogate rollout Root Mean Square Error (RMSE) by 60% (AUV) and 69% (ASV). For the ASV we further demonstrate a Vision language model (VLM)-assisted semantic-mapping pipeline that extracts obstacles and environmental context from satellite imagery, nautical charts, and forecast Application Programming Interface (API) instead of onboard sensors, reaching 96% navigability accuracy as a drop-in replacement for hand-specified obstacle geometry.

自主导航大模型海洋机器人世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。