arXiv:2607.13792cs.CV2026-07被引 1

构建首个面向可穿戴设备的程序性理解问答任务,提升AI对日常动作步骤的推理能力。

EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent

论文配图:EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent
图 1 · 摘自论文原文
  • 设计六类关键步骤问题,评估多模态大模型在第一人称视频中的程序推理能力。
  • 构建含3600个问题的基准数据集,覆盖31项日常任务,揭示现有模型显著不足。
  • 提出自技能探索代理框架,无需标注即可自主发现高效操作策略。

多数日常活动具有程序性特征。然而,现有第一人称视频理解评估大多未涵盖程序性理解,且在主流多模态大模型视频问答(MLLM-VQA)范式下,对复杂关键步骤级推理关注不足。此类能力对可部署于可穿戴设备的程序性AI助手至关重要。为此,我们引入第一人称程序理解问答任务(EgoProceVQA),通过六类以关键步骤为中心的问题,系统评估当前MLLM与智能体的程序推理能力。同时,我们开发了EgoProceGen数据生成平台,高效构建适配不同问题类型的问答数据。基于此平台,我们构建了一个包含3600个问题、4种常见程序场景和31项日常任务的基准数据集。EgoProceVQA的评估显示,现有MLLM与智能体在程序理解方面仍有巨大提升空间。因此,我们进一步提出EgoProceAgent——一种自技能探索智能体框架。该框架设计通用工具库与跨模型共享的标准子技能库,实现无真实标注的自我探索。通过探索子技能的组合与选择方式,智能体发现多样问题的有效技能策略,在多个任务上达到开源模型最优表现。整体上,我们的基准、生成平台与智能体框架共同建立了统一的EgoProceVQA基础。

原文摘要 · Abstract (English)

Most daily activities are inherently procedural. However, existing evaluations for egocentric video understanding seldom address procedural understanding and largely overlook complex key-step-level reasoning under the widely used video question answering (VQA) paradigm for MLLMs. Such capabilities are crucial for building procedural AI assistants deployable on wearable devices. To bridge this gap, we introduce the Egocentric Procedural Understanding VQA task (EgoProceVQA), which systematically evaluates egocentric procedural reasoning abilities of current MLLMs and agents through six types of key-step-centric questions. Furthermore, we develop EgoProceGen, a data generation platform that efficiently constructs QA data tailored to different question types. Based on this platform, we build a benchmark with 3,600 questions, four common procedural scenarios, and 31 everyday procedural tasks. Evaluations on EgoProceVQA show that existing MLLMs and agents still have substantial room for improvement in procedural understanding. Therefore, we further propose EgoProceAgent, a self-skill-exploration agentic framework. We design a generic tool library for procedural understanding and a standardized sub-skill library shared across tools and models, enabling self-exploration without ground-truth supervision. By exploring how to compose and select sub-skills, the agent discovers effective skill strategies for diverse problems, and attains state-of-the-art performance among open-source models on multiple tasks. Together, our benchmark, generation platform, and agentic framework establish a unified foundation for EgoProceVQA. Project page: https://z1oong.github.io/EgoProceVQA/.

程序理解第一人称视频智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。