为助人机器人接触任务设计物理一致评估基准,避免只看完成率的误导。
A Physics-Consistent Benchmark for Contact-Rich Human-Robot Interaction in Assistive Care

- 用可变形假人+物理响应评分,真实模拟人机接触
- 72.9%任务成功率但仅56.4%通过安全筛选,暴露传统评估缺陷
- 适合关注医疗机器人安全性的研究者与开发者
传统任务级评估仅关注机器人是否完成动作,却忽略了物理接触中才显现的失败。这对需要频繁接触的助人任务尤为关键。我们提出一种物理一致的基准,用于接触密集型人机交互,以机器人辅助沐浴为例。该基准结合可变形、被动响应的人体模型、物理感知评分体系以及冻结的视觉/评分员评估协议。通过在医疗假人上进行力-压入测量,校准区域级仿真响应。在冻结的T1-T7协议下,每方法执行140次:基于LLM的状态机达72.9%任务成功率,但通过正确区域和力安全筛选后降至56.4%;VoxPoser产生更轻更稳接触,但仅完成27.9%试验;零样本pi0.5任务成功率仅0.7%,无任何正确区域或安全通过案例。结果表明,任务完成不代表物理合理接触,需在部署前引入物理感知筛选。
原文摘要 · Abstract (English)
Conventional task-level evaluation asks whether a robot policy completes a specified action, but can miss failures that emerge only during physical human contact. This limitation is critical in contact-rich assistive tasks, where meaningful evaluation requires a physically responsive human, interaction-quality assessment beyond task success, and a leak-free observer-scorer protocol. We introduce a physics-consistent benchmark for contact-rich human-robot interaction, instantiated in robot-assisted bathing. The benchmark combines a deformable, passively responding human, physics-aware scores alongside task-level success, and a frozen vision-only / scorer-only evaluation protocol. To establish physical validity, region-wise simulated responses are calibrated against force-indentation measurements from Franka impedance pushes on a medical-care manikin. Under a frozen T1-T7 protocol with 140 runs per method, an LLM-augmented state machine (State Machine) achieves 72.9% task success but drops to 56.4% after correct-region and force-safety screening; VoxPoser produces lighter and more stable contact but completes only 27.9% of trials; and zero-shot pi0.5 achieves 0.7% task success with no correct-region or safety-gated successes. These results show that task completion alone does not imply physically valid contact and motivate physics-aware screening before deployment of contact-rich assistive robot policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。