arXiv:2606.28813cs.ROcs.AI2026-06

从人类视频学动作,让机器人自动适配不同身体和环境。

Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning

论文配图:Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning
图 1 · 摘自论文原文
  • 用物体间交互运动建模人类动作,忽略身体差异
  • 在真实机器人上实现无需实机训练的稳定操作
  • 适合想快速部署机器人抓取与操作的开发者

人类视频是机器人操作任务中可扩展的监督来源,因其数量丰富且自然捕捉了物体间的复杂交互。然而,由于身体形态差异、场景变化及机器人自身可行性约束,将人类示范迁移至机器人仍具挑战。我们提出 Human2Any 框架,通过人类视频学习可复用的以物体为中心的交互先验,无需目标任务场景中的真实机器人演示。Human2Any 以物体-物体交互运动表示操作行为,保留任务相关的场景变化,同时抽象掉与具体身体形态相关细节。该框架将学习到的交互先验与机器人侧的可行性推理及运动规划相结合,使同一套人类知识可适应不同机器人形态、场景几何结构与任务上下文。我们在多种操作场景中验证了 Human2Any 的有效性,包括在 Franka 桌面系统和 RBY-1 人形移动机器人上的真实世界实验,均实现了无需真实机器人训练数据的鲁棒交互式操作。

原文摘要 · Abstract (English)

Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactions. However, transferring human demonstrations to robots remains challenging due to embodiment mismatch, scene variation, and robot-specific feasibility constraints. We present Human2Any, a framework for learning reusable object-centric interaction priors from human videos without requiring real-world robot demonstrations in the target task contexts. Human2Any represents manipulation through object-object interaction motion, capturing task-relevant scene changes while abstracting away embodiment-specific details. It composes learned interaction priors with robot-side feasibility reasoning and motion planning, allowing the same human-derived knowledge to adapt to different embodiments, scene geometries, and task contexts. We validate Human2Any across diverse manipulation settings, including real-world experiments on a Franka tabletop setup and an RBY-1 humanoid mobile robot, demonstrating robust interaction-centric manipulation without real-world robot training data. Project website: https://human2any.github.io/.

机器人操作动作迁移人类示范通用控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。