首个面向第一人称视频流的综合评测基准,检验模型对过去、现在、未来的持续理解能力。
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

- 构建统一流式框架,整合回顾、实时、前瞻三类推理任务
- 涵盖165小时视频与4800个问答对,揭示模型在前瞻性推理上表现不佳
- 发现模型常出现自信错误,适合评估视觉语言模型的时序推理能力
我们提出EgoSAT,首个针对第一人称视频流场景的综合性评测基准,用于评估现代视觉-语言模型(VLMs)在连续视频帧输入下的推理能力。该基准聚焦流式交互理解,要求模型在仅基于已观察帧的前提下,持续解析不断演化的视觉上下文。通过统一多个先前独立的任务,将已完成事件的提问归为回溯推理,正在进行活动的提问视为在线理解,未来行为的提问则涉及前瞻预测。这一统一设定要求模型同时处理过去、现在与未来信息。EgoSAT包含1,997段独特视频,总计约165小时的第一人称影像数据,以及约4,800个高质量问题-答案对,精心设计以探测不同时间上下文中的推理能力。我们在此基准上评估了多种开源与闭源VLM,系统性分析其在流式交互理解中的表现。通过区分可回答性并诊断模型置信度,发现现有模型不仅在前瞻与回溯建模上存在困难,还表现出严重置信度失准:置信度往往无法反映实际可回答性,导致危险的‘自信错误’行为。
原文摘要 · Abstract (English)
We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models (VLMs). The benchmark targets streaming interaction understanding, where video frames arrive sequentially and models must continuously interpret evolving visual context. EgoSAT unifies several previously distinct tasks within a single streaming framework. In this formulation, queries about completed events correspond to retrospective reasoning, queries about ongoing activities require online understanding, and queries about future actions involve prospective anticipation. This unified setting requires models to reason about the past, present, and future while operating under the constraint that only previously observed frames are available. EgoSAT contains 1,997 unique videos spanning 165 hours of egocentric footage and around 4,800 high-quality question-answer pairs, carefully designed to probe reasoning across varying temporal contexts. Using this benchmark, we evaluate a diverse set of both open-weight and closed-weight VLMs, providing a systematic assessment of their ability for streaming interaction understanding. By distinguishing answerability and conducting diagnostics on confidence of models, we find existing models not only struggle with prospective and retrospective modeling, but also exhibit severe mis-calibration: confidence often fails to track inherent answerability, leading to dangerous "confidently wrong" behaviors. Project page: https://leiyj23.github.io/EgoSAT/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。