构建首个面向一周时长第一人称视频的内存推理评测集,挑战模型跨日记忆与推理能力。
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

- 设计三种记忆类型:物体状态、事件顺序和行为模式,覆盖长期视觉理解需求
- 平均每题需回溯25.9小时视频证据,最先进模型准确率仅39.6%
- 揭示跨日记忆仍是未解难题,适合研究长时序多模态系统者参考
下一代视觉助手(如智能眼镜、具身代理、全天候生活记录系统)需对一整天甚至更长时间的连续视觉体验进行推理。在超长视频中,相关信息分散于数小时或数天之间,记忆成为核心挑战:模型必须随时间累积信息、回忆先前状态、追踪时间顺序并抽象重复模式。然而,现有周长视频基准主要针对感知与识别任务(如时刻定位或全局摘要),而非需要跨多日整合证据的推理。为此,我们提出EgoMemReason,一个通过记忆驱动推理评估一周时长第一人称视频理解的综合性基准。该基准涵盖三类互补记忆:实体记忆(跟踪物体状态跨日演变)、事件记忆(回忆并排序相隔数小时或数天的活动)、行为记忆(从稀疏重复观测中抽象出周期性模式)。EgoMemReason包含500个问题,覆盖六项核心挑战,平均每题需5.1段视频证据,平均回溯25.9小时记忆。我们在17种方法(包括多模态大模型与智能体框架)上评估,发现最佳模型整体准确率仅为39.6%。进一步分析表明,三类记忆失败原因各异,且随着证据时间跨度增加,性能持续下降,揭示长时程记忆仍远未解决。我们认为EgoMemReason为评估与推进长上下文、记忆感知的多模态系统奠定了坚实基础。
原文摘要 · Abstract (English)
Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long videos, relevant information is sparsely distributed across hours or days, making memory a fundamental challenge: models must accumulate information over time, recall prior states, track temporal order, and abstract recurring patterns. However, existing week-long video benchmarks are primarily designed for perception and recognition, such as moment localization or global summarization, rather than reasoning that requires integrating evidence across multiple days. To address this gap, we introduce EgoMemReason, a comprehensive benchmark for week-long egocentric video understanding through memory-driven reasoning. EgoMemReason evaluates three complementary memory types: entity memory, tracking how object states evolve and change across days; event memory, recalling and ordering activities separated by hours or days; and behavior memory, abstracting recurring patterns from sparse, repeated observations over the whole week period. EgoMemReason comprises 500 questions across three memory types and six core challenges, with an average of 5.1 video segments of evidence per question and 25.9 hours of memory backtracking. We evaluate EgoMemReason on 17 methods across MLLMs and agentic frameworks, revealing that even the best model achieves only 39.6% overall accuracy. Further analysis shows that the three memory types fail for distinct reasons and that performance degrades as evidence spans longer temporal horizons, revealing that long-horizon memory remains far from solved. We believe EgoMemReason establishes a strong foundation for evaluating and advancing long-context, memory-aware multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。