arXiv:2410.07686cs.RO2024-10被引 12

对比不同输入对无人机强化学习控制的仿真到现实迁移效果

The Power of Input: Benchmarking Zero-Shot Sim-To-Real Transfer of Reinforcement Learning Control Policies for Quadrotor Control

  • 测试多种观测输入配置在仿真环境中的学习表现
  • 发现仅用部分状态信息的策略更易实现零样本跨域迁移
  • 为微型飞行器强化学习输入设计提供实证指导

过去十年中,数据驱动方法因能适应未知或不确定飞行条件而成为无人机控制的热门选择。其中,深度强化学习(DRL)是当前研究最广泛的范式。然而,为微型空中机器人(MAVs)设计DRL智能体仍是开放挑战。尽管已有研究关注智能体的输出配置(即计算何种控制),但对输入数据类型尚未形成共识。许多工作直接向DRL智能体提供完整状态信息,却未质疑这是否冗余、是否过度复杂化学习过程,或是否在真实平台中带来不切实际的信息获取约束。本文对不同观测空间配置进行了深入基准分析。我们在模拟环境中优化多个DRL智能体,采用不同输入选择,并评估其鲁棒性及零样本仿真实现到现实部署的迁移能力。我们相信,本研究基于大量实验结果得出的结论与讨论,将为未来面向飞行机器人任务的DRL智能体研发提供重要里程碑。

原文摘要 · Abstract (English)

In the last decade, data-driven approaches have become popular choices for quadrotor control, thanks to their ability to facilitate the adaptation to unknown or uncertain flight conditions. Among the different data-driven paradigms, Deep Reinforcement Learning (DRL) is currently one of the most explored. However, the design of DRL agents for Micro Aerial Vehicles (MAVs) remains an open challenge. While some works have studied the output configuration of these agents (i.e., what kind of control to compute), there is no general consensus on the type of input data these approaches should employ. Multiple works simply provide the DRL agent with full state information, without questioning if this might be redundant and unnecessarily complicate the learning process, or pose superfluous constraints on the availability of such information in real platforms. In this work, we provide an in-depth benchmark analysis of different configurations of the observation space. We optimize multiple DRL agents in simulated environments with different input choices and study their robustness and their sim-to-real transfer capabilities with zero-shot adaptation. We believe that the outcomes and discussions presented in this work supported by extensive experimental results could be an important milestone in guiding future research on the development of DRL agents for aerial robot tasks.

强化学习无人机控制仿真实现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。