arXiv:2511.20525cs.CV2025-11中稿 · CVPR被引 2

提出细粒度错误归因任务,定位错误的语义、时间与空间位置。

Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos

  • 构建自动标注错误样本的数据引擎,生成大规模带属性注释数据集。
  • 新模型在四项任务上超越已有方法,提升幅度达3%至22%。
  • 适合研究人机交互、视频理解与错误分析的学者使用。

我们提出错误归因(Mistake Attribution, MATT)任务,实现对第一人称视频中人类错误的细粒度理解。以往工作仅检测错误是否发生,而MATT进一步定位错误违反了指令的哪个语义角色、在视频中何时进入不可逆点(点-无回头,PNR),以及在PNR帧中的具体空间位置。我们开发了MisEngine,一个能从现有数据集自动生成丰富属性注释错误样本的数据引擎。将其应用于大规模第一人称视频语料库,得到EPIC-KITCHENS-M和Ego4D-M两个数据集,规模比先前错误数据集大两数量级。随后提出MisFormer,一种统一的基于注意力机制的模型,可同时处理语义、时间与空间维度的错误归因任务,并由MisEngine监督训练。人类评估验证了所生成样本的生态有效性,表明这两个数据集可作为可靠基准。在自建数据集及既有基准上的实验表明,作为单一统一模型,MisFormer在视频-语言理解、时间定位、手-物交互与错误检测四项任务上分别优于当前最优方法至少6.66%、21.81%、18.7%和3.00%。

原文摘要 · Abstract (English)

We introduce Mistake Attribution (MATT), a new task for fine-grained understanding of human mistakes in egocentric videos. While prior work detects whether a mistake occurs, MATT attributes the mistake to what part of the instruction is violated (semantic role), when in the video the deviation becomes irreversible (the Point-of-No-Return, PNR), and where the mistake appears in the PNR frame. We develop MisEngine, a data engine that automatically constructs mistake samples from existing datasets with attribution-rich annotations. Applied to large egocentric corpora, MisEngine yields EPIC-KITCHENS-M and Ego4D-M -- two datasets up to two orders of magnitude larger than prior mistake datasets. We then present MisFormer, a unified attention-based model for mistake attribution across semantic, temporal, and spatial dimensions, trained with MisEngine supervision. A human study demonstrates the ecological validity of our MisEngine-constructed mistake samples, confirming that EPIC-KITCHENS-M and Ego4D-M can serve as reliable benchmarks for mistake understanding. Experiments on both our datasets and prior benchmarks show that MisFormer, as a single unified model, outperforms task-specific SOTA methods by at least 6.66%, 21.81%, 18.7%, and 3.00% in video-language understanding, temporal localization, hand-object interaction, and mistake detection, respectively. Project page: https://yayuanli.github.io/MATT/

错误分析第一人称视频多模态理解细粒度定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。