arXiv:2601.19406cs.ROcs.AI2026-01被引 2

用仿真动作和真人观察数据联合训练,提升机器人操作的数据效率与泛化能力

Sim-and-Human Co-training for Data-Efficient and Generalizable Robotic Manipulation

论文配图:Sim-and-Human Co-training for Data-Efficient and Generalizable Robotic Manipulation
图 1 · 摘自论文原文
  • 融合仿真中的机器人动作先验与真人真实视觉先验
  • 仅用80条真实数据即达62.5%跨域成功率,性能超纯真实数据基线7.1倍
  • 适合数据稀缺但需强泛化的机器人操控任务

合成仿真数据和真实人类数据可有效降低机器人数据采集成本,但前者存在仿真到现实的视觉差距,后者存在人类与机器人身体形态差异。本文发现两者具有天然互补性:仿真提供机器人动作,真人数据提供真实世界观测。基于此,提出SimHum联合训练框架,同时提取仿真中的运动学先验和真人观察中的视觉先验,在真实任务中实现高效且泛化的机器人操控。实验表明,在相同数据预算下,性能最高提升40%;仅使用80条真实数据时,跨域成功率高达62.5%,优于纯真实数据基线7.1倍。

原文摘要 · Abstract (English)

Synthetic simulation data and real-world human data provide scalable alternatives to circumvent the prohibitive costs of robot data collection. However, these sources suffer from the sim-to-real visual gap and the human-to-robot embodiment gap, respectively, which limits the policy's generalization to real-world scenarios. In this work, we identify a natural yet underexplored complementarity between these sources: simulation offers the robot action that human data lacks, while human data provides the real-world observation that simulation struggles to render. Motivated by this insight, we present SimHum, a co-training framework to simultaneously extract kinematic prior from simulated robot actions and visual prior from real-world human observations. Based on the two complementary priors, we achieve data-efficient and generalizable robotic manipulation in real-world tasks. Empirically, SimHum outperforms the baseline by up to $\mathbf{40\%}$ under the same data collection budget, and achieves a $\mathbf{62.5\%}$ OOD success with only 80 real data, outperforming the real only baseline by $7.1\times$. Videos and additional information can be found at \href{https://kaipengfang.github.io/sim-and-human}{project website}.

机器人操控联合训练数据效率泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。