让四足机器人在未知环境中自适应选择最稳健的行走策略。
Risk-Aware Reinforcement Learning with Bandit-Based Adaptation for Quadrupedal Locomotion
- 用风险约束优化训练多套稳健性不同的行走策略。
- 在未知地形中性能接近基线两倍,且两分钟内完成策略选择。
- 无需环境信息,实时根据表现动态调整策略,适合真实场景部署。
本文研究四足机器人行走中的风险感知强化学习。方法通过条件风险价值(CVaR)约束策略优化,训练一组风险可控的策略,提升稳定性和样本效率。部署时,利用仅依赖每轮回报的多臂赌博机框架,从策略族中自适应选择最优策略,无需环境先验信息,可实时应对未知条件。通过在八种未见设置(改变动力学、接触、传感噪声和地形)的仿真中评估,并在Unitree Go2机器人上测试于未见过的地形,结果表明:该风险感知策略在未知环境中平均性能与尾部性能均接近基线的两倍;基于赌博机的自适应机制可在两分钟内完成最佳策略选择。
原文摘要 · Abstract (English)
In this work, we study risk-aware reinforcement learning for quadrupedal locomotion. Our approach trains a family of risk-conditioned policies using a Conditional Value-at-Risk (CVaR) constrained policy optimization technique that provides improved stability and sample efficiency. At deployment, we adaptively select the best performing policy from the family of policies using a multi-armed bandit framework that uses only observed episodic returns, without any privileged environment information, and adapts to unknown conditions on the fly. Hence, we train quadrupedal locomotion policies at various levels of robustness using CVaR and adaptively select the desired level of robustness online to ensure performance in unknown environments. We evaluate our method in simulation across eight unseen settings (by changing dynamics, contacts, sensing noise, and terrain) and on a Unitree Go2 robot in previously unseen terrains. Our risk-aware policy attains nearly twice the mean and tail performance in unseen environments compared to other baselines and our bandit-based adaptation selects the best-performing risk-aware policy in unknown terrain within two minutes of operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。