一个策略通吃多种无人船,无需微调即能精准跟航。
Cross-Platform Control for Autonomous Surface Vehicles via Adaptive Reinforcement Learning

- 用历史交互信息推断船只动力学,实现跨平台自适应控制。
- 真实测试中定位误差比非自适应方法低58%,接近专用调优控制器。
- 仅需简单模型训练,即可零样本部署到不同真实船只上。
自主水面航行器在水动力和推进特性上差异显著,但多数控制器仅针对单一平台设计。本文提出一种自适应强化学习方法,实现轨迹跟踪的零样本跨平台部署,仅用一个策略即可适配不同平台。由于策略无法获知部署平台的动力学特性,我们采用标准的部分可观测性方法,通过教师-学生架构,由学习模块从交互历史中推断出平台动力学的潜在表示。策略在模拟环境中基于随机化船舶动力学进行训练,无需任何微调即可直接部署到两个真实平台。尽管仅使用简单的解析动力学模型而非高保真水动力仿真器,真实实验仍显示该自适应策略相比非自适应学习基线,定位均方误差降低最高达58%,并逼近专用调优控制器的跟踪精度。
原文摘要 · Abstract (English)
Autonomous surface vehicles vary widely in hydrodynamic and actuation characteristics, yet most controllers are designed for single-platform deployment. We present an adaptive reinforcement learning approach for trajectory tracking that enables zero-shot cross-platform deployment using a single policy. Since the deployment platform's dynamics are unknown to the policy, we address cross-platform generalization with the standard partial-observability approach of conditioning on interaction history, employing a teacher-student architecture in which a learned module infers a latent representation of the platform dynamics. The policy is trained in simulation under randomized vessel dynamics and is deployed zero-shot to two real-world platforms without any fine-tuning, despite relying on a simple analytical dynamics model rather than a high-fidelity hydrodynamic simulator. In real-world experiments on two different platforms, the adaptive policy outperforms non-adaptive learning-based baselines by up to 58% in position mean absolute error while approaching the tracking accuracy of a platform-specific tuned controller.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。