arXiv:2608.13113cs.CVcs.AI2026-08

首个月级第一视角视频基准,测试模型长期时空记忆能力。

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

论文配图:EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
图 1 · 摘自论文原文
  • 构建20人持续20-120天的日常影像数据集,覆盖300+小时
  • 模型最高仅71.8%准确率,远低于人类94.2%基准线
  • 揭示现有模型缺乏真实长期记忆,适合研究长时记忆的学者

多模态大语言模型(MLLMs)在视频理解方面取得显著进展,涌现出大量长视频基准。然而,现有基准主要依赖网络来源视频,缺乏片段间的时空连续性,难以评估模型能否维持数日或数周的真实世界经验记忆。本文提出EgoMonth,首个月级第一视角视频理解基准。EgoMonth包含20名参与者持续20至120天的日常生活记录,总计超过300小时,配套1,443个由人工设计的多选问答对。我们设计了一个基于认知的14项评估框架,分为三个层级:情景整合、情景索引与级联推理。对当前最先进的开源与闭源MLLM进行评估发现,表现最佳的Gemini 2.5 Pro仅达71.8%的宏平均准确率,较校正后的人类基线94.2%低22.4个百分点。部分模型在路线推理、跨视角空间推理和方向判断等任务上接近或低于25%随机水平,即使最强闭源模型也显著落后于人类。结果表明,当前MLLM更像有损摘要器而非忠实记忆器,凸显了具备真实长期时空记忆架构的迫切需求。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.

视频理解长时记忆第一视角评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。