arXiv:2606.08102cs.ROcs.AI2026-06

让四足机器人团队持续学习新协作技能,避免遗忘旧技能。

Continual Quadruped Robots Coordination via Semantic Skill Discovery

论文配图:Continual Quadruped Robots Coordination via Semantic Skill Discovery
图 1 · 摘自论文原文
  • 用语义技能库实现任务检索-适配-更新的持续学习流程。
  • 仿真中成功率达95.6%,且无显著灾难性遗忘。
  • 适合需要长期适应新任务的多机器人协作场景。

多四足机器人协作因提升载重能力、扩大接触范围及增强应对复杂任务的适应性而受到关注。现有方法多聚焦于预定义或封闭任务集,常依赖多智能体强化学习训练特定任务的协调策略,但在开放持续学习场景下表现不佳——任务连续到来,机器人需在不遗忘旧技能的前提下习得新协作能力。为此,我们提出Conquer,一个基于语义技能库的持续多四足机器人协调框架,将该问题建模为检索-适配-更新过程。为支持不同规模团队,设计了团队结构化的自盟友目标(SAG)骨干网络,显式建模每个机器人的状态、队友上下文与任务目标。对每项新任务,Conquer利用执行前信息构建任务级语义描述符,并从技能库中检索相关技能进行适配。成功执行后,通过提取轨迹级语义描述符并按语义距离组织,更新技能库,实现持续技能积累与跨任务知识迁移。仿真实验表明,Conquer最终平均成功率可达95.6%,展现出强前向迁移能力与可忽略的灾难性遗忘。真实世界中,基于Unitree Go2团队的演示进一步验证了其部署可行性。模拟与实机演示视频见:https://conquer-project.pages.dev/。

原文摘要 · Abstract (English)

Multi-quadruped coordination has attracted increasing attention due to its enhanced payload capacity, broader contact coverage, and improved adaptability to challenging tasks. Existing methods for multi-quadruped manipulation typically focus on predefined or closed task families, often relying on multi-agent reinforcement learning (MARL) to train task-specific coordination policies. However, such methods struggle in open-ended continual learning settings, where tasks arrive sequentially and robots are expected to acquire new coordination skills while reusing previously learned ones without catastrophic forgetting. To address this challenge, we propose Conquer, a semantic skill-library framework that formulates continual multi-quadruped coordination as a retrieve-adapt-update process. First, to accommodate varying team sizes across tasks, we design a team-structured Self-Allies-Goal (SAG) backbone that supports variable-cardinality robot teams by explicitly modeling each robot's own state, teammate context, and task goal. For each incoming task, Conquer constructs a task-level semantic descriptor from pre-execution information and retrieves a relevant skill from the library for adaptation. After successful execution, Conquer updates the skill library by extracting trajectory-level semantic descriptors and organizing them according to semantic distance, thereby enabling continual skill accumulation and cross-task knowledge transfer. Simulation experiments show that Conquer achieves a final average success rate of 95.6%, demonstrating strong forward transfer and negligible catastrophic forgetting. Real-world rollouts on Unitree Go2 teams further validate the deployment feasibility of Conquer for practical multi-quadruped coordination. Simulation and real-robot demonstration videos are available at: https://conquer-project.pages.dev/.

多机器人持续学习语义技能四足机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。