arXiv:2603.12147cs.CV2026-03

构建细粒度行为意图理解数据集,揭示当前模型依赖静态线索而非时序推理。

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next

  • 设计3014个微观步骤的预结果数据集,标注目标、目的和下一步
  • 模型在无时序信息下表现最优,历史上下文反而降低性能
  • 揭示结果或下一步可提升意图预测,但模型未真正利用过程逻辑

第一人称视频为研究人类行为提供了自然模态,但传统视觉理解主要关注可观测场景、物体和动作,而非组织这些行为的潜在目标。现有意图基准通常聚焦粗粒度事件级目标,忽视意图在操作步骤中的演变。我们提出EgoIntent,一个步骤级意图理解基准,包含来自32段第一人称视频的3,014个预结果微观步骤,覆盖15种室内外日常场景。每个步骤由人工标注三个互补维度:局部意图(What),即即时目标;程序意图(Why),即该步骤在整体流程中的作用;下一步计划(Next),最可能的后续动作。多轮人类评审优化时间边界与标注质量。我们评估了15个跨模态大语言模型,采用基于参考的评分与互补的无参考诊断。对四个代表性模型的受控实验显示,仅一个模型显著受益于正确时序,而单帧边界优于完整有序片段,对三模型表现更佳。仅使用步骤输入对所有四模型最佳,增加15秒历史反而使三模型性能下降。揭示当前结果使局部意图提升7.81分,揭示下一步使下一步预测提升13.17分。这些发现表明,当前模型可通过静态边界线索获得高意图预测分数,但并未真正利用时序顺序或程序历史。

原文摘要 · Abstract (English)

Egocentric video provides a natural modality for studying human behavior, but conventional visual understanding captures mainly observable scenes, objects, and actions rather than the latent goals that organize them. Existing intent benchmarks typically focus on coarse event-level goals and overlook how intent evolves across procedural steps. We introduce EgoIntent, a step-level intent-understanding benchmark comprising 3,014 pre-outcome micro-steps from 32 egocentric videos across 15 indoor and outdoor daily-life scenarios. Each step is manually annotated along three complementary dimensions: Local Intent (What), the immediate goal; Procedural Intent (Why), the role of the step in the broader procedure; and Next-Plan (Next), the action most likely to follow. Multiple rounds of human review refine temporal boundaries and annotation quality. We evaluate 15 multimodal large language models using reference-based scores and complementary reference-free diagnostics. Controlled studies on four representative models show that only one model gains significantly from correct temporal order, while a single boundary frame outperforms the full ordered clip for three models. Step-only input performs best for all four models, and adding 15 seconds of history significantly degrades three. Revealing the current outcome improves Local Intent by 7.81 points, while revealing the following step improves Next-Plan by 13.17 points. These findings indicate that current models can achieve strong intent-prediction scores through static boundary cues without robustly exploiting temporal order or procedural history.

意图理解第一人称视频微步骤分析大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。