构建长时序多源记忆测试集,评估AI对生活习惯与行为推理能力。
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
- 基于认知科学分层结构生成跨时间事件,实现高效可扩展数据合成。
- 顶尖模型在长程记忆任务中仅达55.2%准确率,凸显多源信息融合难度。
- 适用于需理解用户长期行为的个性化智能体研发与评估。
长期记忆是个性化智能体积累知识、推理用户经历并随时间适应的关键。然而现有记忆基准主要针对显式呈现的陈述性记忆(语义与情景记忆),而现实行为还受习惯性与程序性等非陈述性记忆驱动,需从多样数字痕迹中推断。为此,我们提出LifeBench,一个密集连接、长时序事件模拟基准,推动智能体超越简单回忆,实现跨多样化、长时间跨度上下文的陈述性与非陈述性记忆融合推理。构建该基准面临两大挑战:数据质量与可扩展性。我们通过引入真实世界先验(匿名社交调查、地图API、节假日日历)保障数据的真实性、多样性与行为合理性;为提升可扩展性,借鉴认知科学中的部分层级结构,实现高效并行生成的同时保持全局连贯性。性能结果显示,当前最先进记忆系统在该基准上仅达到55.2%准确率,凸显长时序检索与多源整合的固有难度。数据集与生成代码已开源:https://github.com/1754955896/LifeBench。
原文摘要 · Abstract (English)
Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory benchmarks primarily target declarative memory, specifically semantic and episodic types, where all information is explicitly presented in dialogues. In contrast, real-world actions are also governed by non-declarative memory, including habitual and procedural types, and need to be inferred from diverse digital traces. To bridge this gap, we introduce Lifebench, which features densely connected, long-horizon event simulation. It pushes AI agents beyond simple recall, requiring the integration of declarative and non-declarative memory reasoning across diverse and temporally extended contexts. Building such a benchmark presents two key challenges: ensuring data quality and scalability. We maintain data quality by employing real-world priors, including anonymized social surveys, map APIs, and holiday-integrated calendars, thus enforcing fidelity, diversity and behavioral rationality within the dataset. Towards scalability, we draw inspiration from cognitive science and structure events according to their partonomic hierarchy; enabling efficient parallel generation while maintaining global coherence. Performance results show that top-tier, state-of-the-art memory systems reach just 55.2\% accuracy, highlighting the inherent difficulty of long-horizon retrieval and multi-source integration within our proposed benchmark. The dataset and data synthesis code are available at https://github.com/1754955896/LifeBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。