arXiv:2503.07049cs.RO2025-03被引 5

用视觉辅助的师生强化学习,让双足机器人在复杂地形更稳更快

VMTS: Vision-Assisted Teacher-Student Reinforcement Learning for Multi-Terrain Locomotion in Bipedal Robots

  • 构建视觉感知的师生混合专家网络,提升控制适应性
  • 在多种地形上实现稳定行走,性能优于传统模型
  • 适合研究双足机器人运动控制与多地形适应的学者

双足机器人因其类人结构,在诸多场景中具有广泛应用潜力,但其控制受结构复杂性制约。现有研究多依赖本体感知方法,难以应对复杂地形;而视觉感知虽对人机交互环境至关重要,却进一步增加了控制难度。近年来,基于强化学习(RL)的方法在腿式机器人运动控制中展现出前景,尤其在平坦地形表现良好,但在复杂地形适应方面仍存在显著挑战。本文提出一种新型混合专家师生网络强化学习策略,通过简单有效的机制增强基于视觉输入的师生策略性能。该方法结合地形选择策略与教师策略,显著提升学生策略在多样地形中的泛化能力。此外,引入教师-学生网络间的对齐损失,而非强制相似性,进一步优化了学生在不同地形下的导航能力。我们在Limx Dynamic P1双足机器人上进行了实验验证,结果表明该方法在多种地形下均具备可行性与鲁棒性。

原文摘要 · Abstract (English)

Bipedal robots, due to their anthropomorphic design, offer substantial potential across various applications, yet their control is hindered by the complexity of their structure. Currently, most research focuses on proprioception-based methods, which lack the capability to overcome complex terrain. While visual perception is vital for operation in human-centric environments, its integration complicates control further. Recent reinforcement learning (RL) approaches have shown promise in enhancing legged robot locomotion, particularly with proprioception-based methods. However, terrain adaptability, especially for bipedal robots, remains a significant challenge, with most research focusing on flat-terrain scenarios. In this paper, we introduce a novel mixture of experts teacher-student network RL strategy, which enhances the performance of teacher-student policies based on visual inputs through a simple yet effective approach. Our method combines terrain selection strategies with the teacher policy, resulting in superior performance compared to traditional models. Additionally, we introduce an alignment loss between the teacher and student networks, rather than enforcing strict similarity, to improve the student's ability to navigate diverse terrains. We validate our approach experimentally on the Limx Dynamic P1 bipedal robot, demonstrating its feasibility and robustness across multiple terrain types.

双足机器人强化学习视觉感知多地形适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。