构建首个真实佩戴设备流式记忆检索基准,推动可信赖的可穿戴AI记忆系统发展。
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval

- 基于真实佩戴设备视频流,模拟持续记忆检索场景。
- 包含388小时视频与9448个需精确定位的问答对,揭示模型召回率不足人类一半。
- 适用于开发下一代可穿戴AI助手的长期记忆能力评估。
随着可穿戴设备实现持续的第一人称记录,人工智能助手必须具备跨长时程推理以回忆过往经历的能力,即情景记忆。现有基准多依赖可访问完整视频文件的离线评估,无法模拟可穿戴智能的流式现实。我们提出S-EMBER(Streaming Egocentric Memory Benchmark for Episodic Retrieval),一个大规模基准,包含3,141段视频,总计388小时自然活动,由Ray-Ban Meta智能眼镜采集。S-EMBER形式化了基于事实的流式情景记忆检索,从全局离线搜索转向由视觉事件触发的因果、主动回忆。我们提供9,448个问答对,要求通过精确的时间定位进行人工视觉验证,并支持灵活的回答长度,以模拟自然的人机交互。对前沿模型的广泛评测显示存在显著的基于事实召回差距:模型在单独回答和定位时表现尚可,但当两者需同时满足同一查询时,距离人类表现仍有巨大差距,最先进模型的准确率不足人类的一半。S-EMBER为下一代可穿戴AI代理的可靠、基于事实的情景记忆发展奠定了硬件真实的基准。
原文摘要 · Abstract (English)
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory. Current benchmarks often rely on offline evaluation with access to entire video files, failing to simulate the streaming reality of wearable intelligence. We introduce S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses. S-EMBER formalizes grounded streaming episodic retrieval, a paradigm shift from global offline search to causal, active recall triggered by visual events in a continuous stream. We provide 9,448 QA pairs requiring manual visual proof through precise temporal localization and supporting flexible response lengths to simulate natural human-AI interaction. Our extensive benchmarking of frontier models reveals a grounded recall gap: models answer and localize with moderate competence in isolation, yet fall furthest short of human performance when both must hold for the same query, the strongest reaching less than half the human rate. S-EMBER establishes a hardware-authentic foundation for developing grounded, reliable episodic memory in the next generation of wearable AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。