arXiv:2503.13441cs.ROcs.AI2025-03被引 97

用人类第一视角数据训练机器人,提升泛化能力且省时省力。

Humanoid Policy ~ Human Policy

  • 用人类第一视角动作数据统一建模人与机器人的行为
  • 在小规模机器人数据下仍实现良好性能,数据效率显著提升
  • 适合希望降低数据采集成本的机器人研发团队

为多样化任务和平台提升人形机器人操作策略的鲁棒性与泛化能力,需大量数据训练。但仅依赖机器人示范成本高昂,难以扩展。本文探索更易获取的人类第一视角示范作为跨形态训练数据。我们从数据和建模两方面缓解人与人形机器人之间的形态差异:构建了与人形机器人操作对齐的第一视角任务数据集PH2D;提出人类-人形行为策略模型HAT,其状态-动作空间统一,可微分地映射至机器人动作。在少量机器人数据协同训练下,无需额外监督即可直接建模人类与人形机器人作为不同形态。实验证明,使用人类数据能显著提升模型泛化性与鲁棒性,同时大幅提高数据收集效率。

原文摘要 · Abstract (English)

Training manipulation policies for humanoid robots with diverse data enhances their robustness and generalization across tasks and platforms. However, learning solely from robot demonstrations is labor-intensive, requiring expensive tele-operated data collection which is difficult to scale. This paper investigates a more scalable data source, egocentric human demonstrations, to serve as cross-embodiment training data for robot learning. We mitigate the embodiment gap between humanoids and humans from both the data and modeling perspectives. We collect an egocentric task-oriented dataset (PH2D) that is directly aligned with humanoid manipulation demonstrations. We then train a human-humanoid behavior policy, which we term Human Action Transformer (HAT). The state-action space of HAT is unified for both humans and humanoid robots and can be differentiably retargeted to robot actions. Co-trained with smaller-scale robot data, HAT directly models humanoid robots and humans as different embodiments without additional supervision. We show that human data improves both generalization and robustness of HAT with significantly better data collection efficiency. Code and data: https://human-as-robot.github.io/

人形机器人行为迁移数据效率第一视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。