arXiv:2607.13056cs.ROcs.LG2026-07

评测机器人与人协作时的意图理解与协调能力,发现现有模型在互动中表现不佳。

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

论文配图:HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration
图 1 · 摘自论文原文
  • 用可执行的场景脚本建模人机协作中的角色、时间依赖和约束关系。
  • 13个任务超650次测试显示,当前机器人在协调与意图理解上严重不足。
  • 适合研究人机协作、具身智能或交互式机器人系统的人参考。

现有视觉-语言-动作(VLA)基准主要评估孤立的操作技能,而未充分建模人机协作的结构。然而,真实世界协作本质上需要在共享代理下进行协调,包括意图理解、时间同步、协议遵循及动态环境中的安全互动。为填补这一空白,我们提出HRIBench,一个基于可执行交互场景的意图感知人机协作诊断基准。HRIBench将协作任务表示为结构化场景脚本,显式建模代理角色、时间依赖、协调约束和人类行为分布。在此基础上,定义三种代表性交互角色:指导者(Instructor)、合作者(Collaborator)和干扰者(Intruder),涵盖意图沟通、协同操作和人类干预下的鲁棒性。基准包含13个角色条件任务,超过650个评估回合,来自多样化的交互轨迹和场景变化。除二元任务成功外,还引入可解释的以交互为中心的度量指标,涵盖同步性、响应性、协议合规性和安全性。我们在统一协议下评估了基于GR00T、pi0.5和ACT的适配策略。结果表明,尽管具备强操作能力,当前基础机器人策略在协作环境中表现显著不足,暴露出时间协调与意图感知行为的重大缺陷。在HRIBench上微调能持续提升协作性能。在真实世界适应研究中,由HRIBench生成的仿真数据使GR00T N1.5的物理任务成功率从0.10提升至0.43,证明该基准对推动以交互为核心的机器人学习具有重要价值。

原文摘要 · Abstract (English)

Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires coordination under shared agency, including intent understanding, temporal synchronization, protocol adherence, and safe interaction in dynamic environments. To address this gap, we introduce HRIBench, a diagnostic benchmark for intent-aware human-robot collaboration based on executable interaction scenarios. HRIBench represents collaborative tasks as structured scenario scripts that explicitly model agent roles, temporal dependencies, coordination constraints, and human behavior distributions. Building on this abstraction, HRIBench defines three representative interaction roles: Instructor, Collaborator, and Intruder, covering intent communication, joint coordination, and robustness under human intervention. The benchmark contains 13 role-conditioned tasks with over 650 evaluation episodes generated from diverse interaction trajectories and scene variations. Beyond binary task success, HRIBench introduces interpretable interaction-centric metrics spanning synchronization, responsiveness, protocol compliance, and safety. We evaluate adapted policies based on GR00T, pi0.5, and ACT under a unified protocol. Results show that current foundation robot policies struggle substantially in collaborative settings despite strong manipulation ability, revealing major limitations in temporal coordination and intent-aware behavior. Fine-tuning on HRIBench consistently improves collaborative performance. In a real-world adaptation study, simulation data generated by HRIBench improves GR00T N1.5's physical-task success rate from 0.10 to 0.43, demonstrating the benchmark's value for advancing interaction-centric robot learning.

人机协作基准测试意图理解机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。