arXiv:2608.24134cs.CV2026-08

评测视觉模型在第一人称视角下识别操作错误的能力,填补了日常任务理解评估的空白。

EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI

论文配图:EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI
图 1 · 摘自论文原文
  • 提出基于第一人称视频的程序性错误问答任务,显式建模操作错误。
  • 发现现有模型在识别不同错误类型上普遍存在不足,尤其对序列依赖错误敏感度低。
  • 引入自适应解耦推理框架,显著提升模型对复杂操作错误的理解能力,适合辅助机器人研究者使用。

我们日常生活中的大多数活动都是程序性的,由一系列相互依赖的步骤构成。然而,现有的视觉智能体与视觉语言模型(VLMs)基准测试普遍忽略了从第一人称视觉角度评估其程序性理解能力,特别是检测程序性错误这一关键能力。为此,本文首次提出EgoErrorVQA任务,用于第一人称程序性理解的评估,并显式建模程序性错误。同时,我们基于Agent2Agent(A2A)协议开发了一个用户友好的评估代理,通过基于VQA的交互实现对视觉智能体的严谨、标准化评估。在开放问答和多选题两种形式下对多种模型进行评估,揭示了其在处理程序性错误及错误类型方面的持续薄弱。此外,我们提出Ego-ADR——一种自适应解耦推理框架,通过解耦复杂的程序推理过程来增强模型对程序性错误的理解。该框架在多个指标上超越选定基线,达到可比设置下的最先进水平。

原文摘要 · Abstract (English)

The majority of our everyday activities are procedural and consist of sequences of interdependent steps. However, existing benchmarks for Visual Agents and Visual Language Models (VLMs) overlook the evaluation of their procedural comprehension ability from an egocentric visual perspective, particularly for detecting procedural errors, a critical capability for everyday assistance. To bridge this gap, the EgoErrorVQA task is firstly proposed for egocentric procedural comprehension with explicit procedural errors modeling. Besides, we develop a user-friendly evaluator agent based on the Agent2Agent (A2A) protocol, enabling rigorous and standardized evaluation of visual agents through VQA-based interaction. A range of models are evaluated using both open-ended and multiple-choice questions, revealing persistent weaknesses in handling procedural errors and error types. Moreover, we introduce Ego-ADR, an Adaptive Decoupled Reasoning framework that decouples complex procedural reasoning to enhance models' understanding of procedural errors. It achieves consistent performance gains over the selected baselines and attains state-of-the-art results on several metrics under comparable settings. Code: https://github.com/z1oong/EgoErrorVQA

第一人称视觉程序性理解错误检测视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。