arXiv:2509.18631cs.ROcs.AI2025-09NeurIPS被引 13

用少量真实数据提升仿真训练的机器人抓取能力

Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training

  • 联合优化仿真与真实数据,通过最优传输对齐观测与动作联合分布
  • 在真实场景中成功率提升30%,且能泛化到仅在仿真中见过的场景
  • 适合希望低成本获取强泛化能力机器人策略的研究者

行为克隆在机器人操作中展现潜力,但大规模获取真实世界示范成本高昂。尽管模拟数据可提供可扩展替代方案,尤其借助自动化示范生成技术,但将策略迁移到真实世界仍受仿真与真实域间差异的阻碍。本文提出一种统一的仿真-真实协同训练框架,旨在学习具备泛化能力的操作策略,主要依赖仿真数据,仅需少量真实示范。核心在于学习一个域不变、任务相关的特征空间。关键洞见是:跨域对观测及其对应动作的联合分布进行对齐,比单独对齐观测(边缘分布)提供更丰富的信号。我们通过在协同训练框架中嵌入受最优传输启发的损失实现这一目标,并进一步扩展为非平衡最优传输框架,以应对大量仿真数据与有限真实样本之间的不平衡。我们在具有挑战性的操作任务上验证了该方法,结果表明其可利用丰富的仿真数据,使真实世界成功率提升高达30%,并能泛化至仅在仿真中出现的场景。项目网页:https://ot-sim2real.github.io/。

原文摘要 · Abstract (English)

Behavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonstration generation, transferring policies to the real world is hampered by various simulation and real domain gaps. In this work, we propose a unified sim-and-real co-training framework for learning generalizable manipulation policies that primarily leverages simulation and only requires a few real-world demonstrations. Central to our approach is learning a domain-invariant, task-relevant feature space. Our key insight is that aligning the joint distributions of observations and their corresponding actions across domains provides a richer signal than aligning observations (marginals) alone. We achieve this by embedding an Optimal Transport (OT)-inspired loss within the co-training framework, and extend this to an Unbalanced OT framework to handle the imbalance between abundant simulation data and limited real-world examples. We validate our method on challenging manipulation tasks, showing it can leverage abundant simulation data to achieve up to a 30% improvement in the real-world success rate and even generalize to scenarios seen only in simulation. Project webpage: https://ot-sim2real.github.io/.

机器人操控域适应协同训练最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。