arXiv:2609.07445cs.RO2026-09

用物理仿真训练多艘拖船协同推船,提升海上作业的精准与适应性。

SMaRT-Tug: Structured Multi-Agent Reinforcement Learning for Physics-Based Tugboat-Barge Collaborative Manipulation

论文配图:SMaRT-Tug: Structured Multi-Agent Reinforcement Learning for Physics-Based Tugboat-Barge Collaborative Manipulation
图 1 · 摘自论文原文
  • 基于物理引擎构建多智能体强化学习框架,支持真实海洋环境模拟。
  • 在直行、转向和减速任务中表现优于传统控制器和集中式算法。
  • 零样本泛化到更复杂海况和更多拖船,训练仅用两艘却可扩展至四艘。

自主拖船系统是港口物流与船舶操纵自动化的核心,需多艘拖船协同运输大型船只。此类协作推拽面临流体耦合、低阻力、强环境干扰、欠驱动的船体动力学及频繁接触交互等挑战。传统控制依赖简化模型与固定配置,适应性差;学习方法受限于缺乏可扩展且物理真实的训练环境。本文提出一个基于物理的GPU加速仿真与学习框架,集成定制浮力模型、波浪建模与水动力阻力,支持大规模多智能体训练。在该环境中,采用带结构化控制先验(SCP)的去中心化MAPPO策略,提升训练稳定性并维持合理推拽构型。评估显示,该框架在直线航行、转向与减速任务中优于基于PID的控制器和集中式PPO基线。进一步验证了对更复杂海况和高级操作的零样本泛化能力,以及从两艘拖船扩展至三艘、四艘的零样本可扩展性。

原文摘要 · Abstract (English)

Autonomous tugboating is central for automating maritime operations such as port logistics and vessel maneuvering, where multiple tugboats must cooperatively transport/manipulate a larger vessel. Collaborative pushing in this setting is challenging due to coupled hydrodynamics, low resistance, strong environmental disturbances, underactuated barge dynamics, and contact-rich interactions. Conventional control methods often rely on simplified models and fixed configurations, which limit their adaptability, while learning-based approaches are constrained by the lack of scalable and physically realistic training environments. We address these challenges by introducing a physics-based, GPU-accelerated simulation and learning framework for collaborative tugboat manipulation. Our simulator incorporates a customized buoyancy model, wave modeling, and hydrodynamic resistance, and supports large-scale multi-agent training under marine dynamics. In this simulator, we train a decentralized MAPPO (Multi-Agent PPO) policy augmented with a structured control prior (SCP) to improve training stability and maintain feasible pushing configurations. We evaluate our learned policy on straight-line transit, turning, and deceleration tasks, where we show that our decentralized framework yields more reliable and accurate maneuvering performance compared to a PID-based controller and a centralized PPO baseline. We further demonstrate zero-shot generalization to more challenging sea states and advanced maneuvers, as well as zero-shot scalability to larger teams of three and four tugboats despite training with only two agents.

多智能体强化学习物理仿真海上操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。