将强化学习融入传统机器人架构,实现足球机器人胜出
Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer
- 在经典机器人框架中嵌入强化学习,分层控制行为决策
- 通过多保真度仿真到真实环境迁移,实现2024年机器人足球赛胜利
- 适合研究机器人自主决策与仿真到现实迁移的开发者
部分可观测、实时、动态且多智能体环境中的机器人决策仍是未解难题。无模型强化学习(RL)虽具潜力,但在复杂环境中端到端训练常不可行。为应对RoboCup标准平台联赛(SPL)中的挑战,我们提出一种新架构:将强化学习集成于经典机器人栈,采用多保真度仿真到现实迁移策略,并将行为分解为可学习的子行为,由启发式选择机制协调。该系统在2024年RoboCup SPL挑战赛盾牌组中取得胜利。本文详述系统架构并实证分析关键设计决策对成功的影响。结果表明,基于强化学习的行为可有效融入完整机器人行为体系。
原文摘要 · Abstract (English)
Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) is a promising approach to learning decision-making in such domains, however, end-to-end RL in complex environments is often intractable. To address this challenge in the RoboCup Standard Platform League (SPL) domain, we developed a novel architecture integrating RL within a classical robotics stack, while employing a multi-fidelity sim2real approach and decomposing behavior into learned sub-behaviors with heuristic selection. Our architecture led to victory in the 2024 RoboCup SPL Challenge Shield Division. In this work, we fully describe our system's architecture and empirically analyze key design decisions that contributed to its success. Our approach demonstrates how RL-based behaviors can be integrated into complete robot behavior architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。