arXiv:2605.13665cs.RO2026-05

用强化学习让四足机器人学会在复杂隧道中灵活行走。

Robot Squid Game: Quadrupedal Locomotion for Traversing Narrow Tunnels

论文配图:Robot Squid Game: Quadrupedal Locomotion for Traversing Narrow Tunnels
图 1 · 摘自论文原文
  • 通过程序生成隧道环境,训练专家策略后蒸馏知识到统一策略
  • 无需复杂奖励设计,实现在多种狭窄隧道中稳定通行
  • 适合需要在废墟、洞穴等狭小空间作业的机器人应用

四足机器人在搜救和基础设施巡检等关键任务中展现出巨大潜力,但自主穿越隧道、洞穴和坍塌结构等受限三维环境仍是重大挑战。现有方法常受限于僵化的步态模式、对多样几何结构适应性差,以及对环境假设过于简化。本文提出一种结合过程化环境生成与策略蒸馏的强化学习框架。采用教师-学生训练范式,先在程序生成的隧道几何中训练多个专家策略,再将导航能力蒸馏至统一的学生策略。该方法避免了端到端强化学习中的复杂奖励设计,将复杂任务分解为更易学习的子任务。通过训练时合成多样化隧道结构并提炼通用导航策略,该方法在多种复杂空间约束下实现稳定通行,传统方法失效的场景中表现优异。仿真与真实实验均验证了其有效性。

原文摘要 · Abstract (English)

Quadruped robots demonstrate exceptional potential for navigating complex terrain in critical applications such as search and rescue missions and infrastructure inspection However autonomous traversal of confined 3D environments including tunnels caves and collapsed structures remains a significant challenge Existing methods often struggle with rigid gait patterns limited adaptability to diverse geometries and reliance on oversimplified environmental assumptions This paper introduces a Reinforcement Learning RL framework that combines procedural environment generation with policy distillation to enable robust locomotion across various tunnel configurations Our approach leverages a teacher student training paradigm where specialized expert policies trained on procedurally generated tunnel geometries transfer their knowledge to a unified student policy This strategy eliminates the need for complex reward shaping in end-to-end RL training simplifying the process by breaking down complicated tasks into smaller more manageable components that are easier for the robot to learn By synthesizing diverse tunnel structures during training and distilling navigation strategies into a generalizable policy our method achieves consistent traversal across complex spatial constraints where conventional approaches fail We demonstrate through both simulation and real world experiments that our method enables quadruped robots to successfully traverse challenging confined tunnel environments

四足机器人强化学习隧道行走策略蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。