arXiv:2412.17730cs.ROcs.CV2024-12被引 23

首个支持通用人形机器人场景交互学习的模仿基准,基于大规模人类动作数据。

Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking

  • 用大规模合成与真实人类动作数据构建通用模仿学习基准。
  • 涵盖11000种物体形状和23000条动作参考,覆盖6类家庭场景任务。
  • 适合研究机器人运动重定向、模仿学习及跨场景泛化方向的学者。

通过模仿人类数据学习人形机器人在三维场景中通用交互技能,是机器人领域的重要挑战,具有广泛的实际应用价值。然而,现有方法与评测基准受限于小规模、人工收集的示范数据,缺乏支持场景几何泛化的大型数据集与评测体系。为此,我们提出 Mimicking-Bench,首个面向通用人形-场景交互学习的模仿基准,基于大规模人类动画参考构建。该基准包含6类家庭环境中的全身人形-场景交互任务,覆盖11,000种多样物体形状,以及20,000条合成和3,000条真实世界的人类交互技能参考。我们构建了完整的技能学习流水线,涵盖运动重定向、运动追踪、模仿学习及其组合,并开展全面实验,验证了人类模仿在技能学习中的价值,揭示关键挑战与未来研究方向。

原文摘要 · Abstract (English)

Learning generic skills for humanoid robots interacting with 3D scenes by mimicking human data is a key research challenge with significant implications for robotics and real-world applications. However, existing methodologies and benchmarks are constrained by the use of small-scale, manually collected demonstrations, lacking the general dataset and benchmark support necessary to explore scene geometry generalization effectively. To address this gap, we introduce Mimicking-Bench, the first comprehensive benchmark designed for generalizable humanoid-scene interaction learning through mimicking large-scale human animation references. Mimicking-Bench includes six household full-body humanoid-scene interaction tasks, covering 11K diverse object shapes, along with 20K synthetic and 3K real-world human interaction skill references. We construct a complete humanoid skill learning pipeline and benchmark approaches for motion retargeting, motion tracking, imitation learning, and their various combinations. Extensive experiments highlight the value of human mimicking for skill learning, revealing key challenges and research directions.

人形机器人模仿学习场景泛化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。