arXiv:2605.19559cs.CVcs.AI2026-05

构建首个可验证的沉浸式视频操作推理基准,评估模型是否真懂动作逻辑。

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

论文配图:EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
图 1 · 摘自论文原文
  • 基于时空图生成细粒度推理链,标注每一步理由与证据
  • 包含3172个可验证问答对,覆盖感知、预测与高层推理任务
  • 发现多数模型解释看似正确,实则证据与答案不符,适合研究可解释性

多模态大模型在第一人称视频理解方面发展迅速,尤其关注精细的手物交互识别、物体状态时序追踪以及动态环境下的操作过程推理。然而,现有第一人称视频基准普遍存在**缺乏可落地的推理评估**问题,难以支持细粒度的操作中心型推理,且极少检验模型推理是否基于明确的时空证据。为填补此空白,本文提出**EgoCoT-Bench**,一个用于可验证、有依据的操作中心型推理的第一人称视频基准,包含显式的分步推理标注。该基准涵盖351段第一人称视频,共3,172个可验证问答对,分为四大任务组和12个子任务组,涵盖感知与回溯、预测及高层推理。数据通过时空场景图(STSG)引导生成,并经人工精修以确保准确性、第一人称相关性与细粒度质量。实验表明,模型在第一人称细粒度推理仍面临持续挑战,且许多多模态模型虽答案正确,但其证据与答案不一致。我们希望EgoCoT-Bench能成为第一人称视频理解中可验证推理的重要测试平台。项目页面与补充材料见:https://dstardust.github.io/EgoCoT/。

原文摘要 · Abstract (English)

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state changes over time, and reason about manipulative processes in dynamic environments from a first-person perspective. However, existing egocentric video benchmarks suffer from \textbf{limited grounded rationale evaluation}, offering limited support for fine-grained operation-centric reasoning and rarely examining whether model rationales are grounded in explicit spatio-temporal evidence. To address this gap, we introduce \textbf{EgoCoT-Bench}, a fine-grained egocentric benchmark for grounded and verifiable operation-centric reasoning with explicit step-by-step rationale annotations. Overall, EgoCoT-Bench comprises 3,172 verifiable QA pairs over 351 egocentric videos separated into four task groups for a total of 12 sub-task groups, encompassing perception and retrospection, anticipation, and high-level reasoning. The benchmark is constructed through a spatio-temporal scene graphs (STSG) guided generation framework and is further refined by human annotators to ensure correctness, egocentric relevance and fine-grained quality. Experimental results show continuing difficulties with egocentric fine-grained reasoning and further reveal that many multimodal models produce explanations that are answer-correct, but have evidence that is inconsistent with the answer. We hope EgoCoT-Bench can serve as a useful testbed for grounded and verifiable reasoning in egocentric video understanding. Project page and supplementary materials are available at: https://dstardust.github.io/EgoCoT/.

第一人称视频可验证推理多模态模型操作链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。