arXiv:2607.29433cs.CL2026-07

发现大模型记得用户偏好却不用,揭示记忆利用的深层瓶颈

Know It, Act on It: Investigating Memory Utilization in LLM Personalization

论文配图:Know It, Act on It: Investigating Memory Utilization in LLM Personalization
图 1 · 摘自论文原文
  • 设计分离测试:先考记忆再考应用,精准定位失效环节
  • 1000条偏好测试显示,记住了但用不上的比例超六成
  • 健康类偏好最难落地,对真实场景影响最大

随着大语言模型代理演变为个性化伙伴,记忆能力日益关键。然而,这些模型存在知识利用难题:即使用户偏好已完整出现在上下文中,仍可能无法据此调整行为。当代理在应体现用户偏好的情境中表现失准时,难以判断是忘记信息,还是记住却未使用。为厘清这一问题,我们提出解耦评估范式,对同一用户偏好同时实施‘知道’与‘行动’测试。我们在16个系统和5种记忆架构上开展大规模实验,评估了1000条嵌入在三种表达强度下的用户偏好。结果表明,‘知道’与‘行动’之间存在显著差距:多数代理能通过记忆测试,但在对应行为场景中未能体现该偏好。尽管不同记忆架构有助于缩小差距,但在健康与心理治疗相关偏好上,利用效果依然薄弱,而这些领域的执行失败具有最高现实风险。

原文摘要 · Abstract (English)

As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they are fully present in context. When an agent fails to tailor its response in a context where previously shared user preferences should matter, it is unclear whether the model failed to remember that information or remembered it but failed to use it. To isolate this breakdown, we introduce a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference. We conduct large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength. Our results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioral scenario. While memory architectures reduce this gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes.

大模型个性化记忆机制行为对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。