arXiv:2602.18025cs.AIcs.RO2026-02被引 1

用跨体态离线强化学习,让不同机器人的数据协同训练,提升泛化能力。

Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets

  • 将异构机器人数据融合,通过离线强化学习统一训练策略。
  • 在含大量低质量轨迹的数据集上表现优于行为克隆,尤其在16种机器人平台上。
  • 按形态分组减少冲突,简单分组策略优于复杂冲突缓解方法。

大规模机器人策略预训练受限于每种平台高质量示范数据的高成本采集。本文结合离线强化学习与跨体态学习,利用专家与大量次优数据,并整合多种形态机器人的轨迹以获得通用控制先验。我们构建了涵盖16种不同机器人平台的运动数据集,系统分析该范式的优势与局限。实验表明,该方法在富含次优轨迹的数据集上预训练效果优于纯行为克隆;但随着次优数据比例和机器人类型增加,跨形态梯度冲突开始阻碍学习。为此,我们提出基于体态相似性的静态分组策略,将机器人按形态聚类后以群体梯度更新模型,显著降低跨机器人冲突,性能超越现有冲突缓解方法。

原文摘要 · Abstract (English)

Scalable robot policy pre-training has been hindered by the high cost of collecting high-quality demonstrations for each platform. In this study, we address this issue by uniting offline reinforcement learning (offline RL) with cross-embodiment learning. Offline RL leverages both expert and abundant suboptimal data, and cross-embodiment learning aggregates heterogeneous robot trajectories across diverse morphologies to acquire universal control priors. We perform a systematic analysis of this offline RL and cross-embodiment paradigm, providing a principled understanding of its strengths and limitations. To evaluate this offline RL and cross-embodiment paradigm, we construct a suite of locomotion datasets spanning 16 distinct robot platforms. Our experiments confirm that this combined approach excels at pre-training with datasets rich in suboptimal trajectories, outperforming pure behavior cloning. However, as the proportion of suboptimal data and the number of robot types increase, we observe that conflicting gradients across morphologies begin to impede learning. To mitigate this, we introduce an embodiment-based grouping strategy in which robots are clustered by morphological similarity and the model is updated with a group gradient. This simple, static grouping substantially reduces inter-robot conflicts and outperforms existing conflict-resolution methods.

离线强化学习多机器人跨体态策略预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。