构建机器人故障感知与恢复的基准数据集,助力更智能的人机交互系统。
REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery

- 基于214次交互实验,融合多模态数据建模用户对故障的长期适应行为。
- 提出三类新任务:跨会话故障检测、细粒度故障分类、用户偏好恢复预测。
- 适合作为医疗人机交互研究及鲁棒机器人系统开发的评估工具。
理解用户对机器人故障的感知与响应对于构建稳健可信的机器人系统至关重要。以往工作通常将故障视为独立事件,强调二元故障检测,并采用规则化恢复建模。本文提出REPAIR-Bench,基于41名参与者完成的214次交互实验,涵盖四种人为诱导的故障类型,提供同步的面部动作单元、头部姿态、语音转录文本,以及交互后的主观情感与恢复报告。该基准包含三项创新评估任务,共同刻画人机交互中故障的全生命周期:(i) 跨依赖交互会话的故障检测,建模用户在重复故障下的长期适应;(ii) 超越二元成败的视觉故障类型分类;(iii) 用户中心的恢复策略预测,从交互上下文推断用户偏好的恢复方式,而非依赖人工设计或规则。基线实验显示,层次化递归模型在故障检测上优于单会话模型(严格F1: 0.80 vs. 0.68),故障定位均值偏差-0.51秒,中位绝对误差2.97秒;恢复预测方面,经QLoRA微调的Mistral-7B达到Hit@5=0.76,F1@5=0.32。REPAIR-Bench为HRI及医疗人机交互领域提供了标准化框架,用于评估机器人故障并构建透明、自适应、可信的恢复系统。
原文摘要 · Abstract (English)
Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i) often treats failures as independent events, (ii) emphasizes binary failure detection, (iii) with rule-based recovery modeling. We present REPAIR-Bench, built on 214 interaction trials from 41 participants, the benchmark spans four induced failure types and provides synchronized facial action units, head pose, speech transcripts, and post-interaction affect and recovery reports. The benchmark spans three novel evaluation tasks that jointly capture the lifecycle of failure in human-robot interaction (HRI): (i) failure detection over inter-dependent interaction sessions, modeling longitudinal user adaptation across repeated failures; (ii) visual failure-type classification beyond binary success/failure formulations; and (iii) user-centered recovery prediction, inferring users' preferred recovery strategies from interaction context rather than relying on manually designed or rule-based strategies. In baseline experiments, hierarchical recurrent modeling improved failure detection over a single-session model (strict F1: 0.80 vs. 0.68), achieved a failure localization mean signed error of -0.51 s, median absolute error of 2.97 s and, for recovery prediction, a QLoRA-tuned Mistral-7B reached Hit@5=0.76 and F1@5=0.32. REPAIR-Bench provides both the HRI and Medical HRI communities with a standardized framework for (1) evaluating robot failures and (2) building transparent, adaptive, and trustworthy recovery systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。