arXiv:2608.13049cs.ROcs.CV2026-08被引 1

构建人类到机器人操作视频生成的评测基准,验证模型跨体感迁移能力

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

论文配图:H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
图 1 · 摘自论文原文
  • 提出H2R-Bench基准,评估视频世界模型将人类视角操作视频转为机器人动作的能力
  • 11个主流模型在6类操作任务中表现有限,普遍存在体感一致性与功能交互失败
  • 适合研究机器人学习、跨体感迁移和视频生成的学者使用

大规模操作数据对机器人学习至关重要,但收集机器人示范成本高且难扩展。相比之下,大量第一人称人类操作视频蕴含丰富行为经验,但因人类手部与机器人末端执行器差异,跨体感迁移仍具挑战。近期视频世界模型为从人类观察合成机器人中心操作视频提供了新路径,但其跨体感迁移能力尚未被系统评估。为此,我们提出H2R-Bench,一个用于评估人类到机器人操作视频生成的基准,要求模型在指定体感约束下,将第一人称人类示范视频转换为机器人操作视频。每个测试实例包含人类示范视频、目标体感约束及源基标注(涵盖任务目标、动作事件、功能接触与物体响应)。该基准从五个维度评估生成视频:目标状态完成度、动作事件完成度、功能接触转移、体感正确性与整体视频质量。我们在六类操作任务和两种机器人体感上评测了11个前沿视频生成模型。结果显示,当前视频世界模型在人类到机器人的操作迁移中仍存在明显局限:即使领先模型也常出现体感不一致、功能交互失效和任务执行失败等问题。H2R-Bench为系统诊断视频世界模型是否能弥合人类与机器人体感差距,并将人类操作观察转化为机器人训练资源提供了框架。

原文摘要 · Abstract (English)

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

视频生成机器人学习跨体感迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。