arXiv:2606.23153cs.RO2026-06

利用非对称物理模拟,让数百只四足机器人高效学习协同避障。

Asymmetric physics enables efficient learning in quadrupedal robot swarms

论文配图:Asymmetric physics enables efficient learning in quadrupedal robot swarms
图 1 · 摘自论文原文
  • 分层训练:真实物理仿真与可微代理模型结合,提升学习效率。
  • 成功训练512只机器人在复杂环境中协同导航,零样本迁移至6台实体机器人。
  • 无需通信或全局地图,适用于森林、迷宫等多场景,具备自然避让行为。

动物群体通过局部协作在复杂环境中导航,但机器人集群在现实世界中仍难以实现这一能力。端到端学习为协调控制提供了路径,但将其扩展到具身集群时面临挑战:当视觉感知、密集的机器人-机器人交互和接触丰富的运动需共同学习时,基于采样的强化学习效率低下。本文提出,通过非对称物理建模,可实现大规模四足机器人集群的视觉驱动、去中心化控制的高效端到端学习。训练中,四足机器人在共享环境中互动,高保真、不可微的仿真器生成真实的运动与接触动力学,而可微代理模型则提供导航与运动策略的梯度。该分离机制使多达512个机器人能在障碍物密集环境中学习协调导航策略。部署时,每台机器人仅依赖单个前向深度相机,无需显式通信、集中规划或全局地图。策略在森林、桥梁、围栏、狭窄通道和迷宫等场景中具有泛化能力,并实现零样本迁移至六台实体机器人,覆盖五个真实场景。集群表现出预测性避让、靠右通行、瓶颈处暂停、沿墙行走等行为,证明非对称物理建模可高效训练可扩展的四足机器人集群去中心化控制策略。

原文摘要 · Abstract (English)

Animal collectives navigate cluttered environments through local coordination, yet robot swarms still struggle to reproduce this capability in the physical world. End-to-end learning offers a route to such coordination, but scaling it to embodied swarms remains difficult: standard sampling-based reinforcement learning becomes inefficient when visual perception, dense robot-robot interaction, and contact-rich locomotion must be learned together. Here we show that asymmetric physics enables efficient end-to-end learning of vision-based, decentralized control in large swarms of quadrupedal robots. During training, quadrupeds interact in shared environments, where a high-fidelity, non-differentiable simulator generates realistic motion and contact dynamics, and differentiable surrogate models provide gradients for navigation and locomotion policies. This separation enables up to 512 quadrupeds to learn coordinated navigation policies in obstacle-rich environments. At deployment, each robot acts from a single forward-facing depth camera, without explicit communication, centralized planning, or global maps. The policies generalize across forests, bridges, enclosures, narrow passages, and mazes, and zero-shot transfer to six physical quadrupeds across five real-world scenarios. The resulting swarms exhibit predictive avoidance, right-side yielding, pausing before bottlenecks, and wall following, showing that asymmetric physics enables efficient training of scalable decentralized control policies for quadrupedal robot swarms.

四足机器人集群智能强化学习物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。