用自传式叙事构建长期记忆评估基准,更真实测试对人的理解能力。
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
- 基于自传体叙述构建带时间锚点的回溯式数据流
- 检索增强模型仅提升事实准确率,高阶推理仍存错误
- 适合研究长期记忆与人格理解的AI系统评估
现有长时程记忆评测多依赖多轮对话或合成用户历史,导致检索性能无法准确反映对人的理解。我们提出 extit{KnowMe-Bench},一个公开可获取的基准,数据源自长篇自传式叙述,包含行为、上下文与内心想法,为推断稳定动机与决策原则提供密集证据。该基准将每段叙述重构为带回溯感知、时间锚定的数据流,并通过关联证据的问题评估模型在事实回忆、主观状态归因与原则级推理上的表现。在多种叙述来源中,检索增强系统主要提升事实准确性,但在时间锚定解释和更高层次推理上仍存在错误,凸显了超越检索的记忆机制的必要性。数据集已开源:https://github.com/QuantaAlpha/KnowMeBench。
原文摘要 · Abstract (English)
Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding. We present \BenchName, a publicly releasable benchmark built from long-form autobiographical narratives, where actions, context, and inner thoughts provide dense evidence for inferring stable motivations and decision principles. \BenchName~reconstructs each narrative into a flashback-aware, time-anchored stream and evaluates models with evidence-linked questions spanning factual recall, subjective state attribution, and principle-level reasoning. Across diverse narrative sources, retrieval-augmented systems mainly improve factual accuracy, while errors persist on temporally grounded explanations and higher-level inferences, highlighting the need for memory mechanisms beyond retrieval. Our data is in \href{KnowMeBench}{https://github.com/QuantaAlpha/KnowMeBench}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。