arXiv:2607.11570cs.ROcs.HC2026-07

构建真实场景下人机交互的多模态错误检测与预判数据集

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions

  • 基于自然场景视频,构建双数据集支持多模态建模
  • 三支团队模型均超越基准,实现误差检测与预判
  • 适合关注人机共融、行为预测的研究者

随着机器人深入人类环境,其检测与响应错误的能力对维持用户信任和交互质量至关重要。尽管机器学习取得进展,现有方法多局限于特定场景、受控环境或预提取特征,泛化能力有限。为此,第三届ERR@HRI挑战赛(ERR@HRI 3.0)提供两个互补数据集,支持从头到尾的端到端创新:(1) 观众情绪检测(BAD)数据集,包含45名参与者在机器人与人类失败场景下的自发反应视频;(2) 错误预判(Bad Idea)数据集,记录29名参与者在故障发生前预测动作结果时的预期面部反应。两数据集通过众包采集,捕捉真实世界中的自然变异性。参赛者针对旁观者反应检测(任务1)、前瞻性结果预测(任务2)开发多模态模型,另有跨数据集泛化任务(任务3)。三支队伍提交有效模型,均优于卷积神经网络基线。本文详述数据集、任务、基线与结果,并讨论构建可泛化、情境感知且具备预判能力的人机交互错误检测系统的意义。

原文摘要 · Abstract (English)

As robots become increasingly integrated into human environments, their ability to detect and respond to errors remains critical for maintaining user trust and interaction quality. While recent advances in machine learning have improved error detection capabilities, most approaches are limited to specific contexts, controlled settings, or pre-extracted features, limiting their generalizability and applicability to real-world conditions. To address this challenge, the third edition of the ERR@HRI Challenge (ERR@HRI 3.0) provided researchers with two complementary datasets that enable end-to-end innovation in methods for both detecting and preventing errors in human-robot interaction. The challenge offered raw, non-anonymized video data from naturalistic settings: (1) the Bystander Affect Detection (BAD) dataset, containing webcam recordings of 45 participants' spontaneous reactions to robot and human failure scenarios; and (2) the Bad Idea dataset, featuring 29 participants' anticipatory facial responses while predicting action outcomes before failures occur. Both datasets were collected via crowdsourcing, capturing the inherent variability of real-world conditions. This naturalistic variability, while challenging, provides an authentic testbed for developing robust error detection systems. Participants developed multimodal machine learning models for bystander reaction detection (Track 1) and anticipatory outcome prediction (Track 2), with an optional cross-dataset generalization track (Track 3). Three teams submitted valid models, all of which surpassed our convolutional neural network baselines. This paper describes the datasets, tasks, baselines, and results of ERR@HRI 3.0, and discusses implications for building generalizable, context-aware, and anticipatory error detection systems for human-robot interaction.

人机交互多模态错误检测预判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。