arXiv:2603.25973cs.CL2026-03被引 15

首个基于真实用户行为的长时记忆评估基准,测试大模型跨领域长期个性化能力。

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization

  • 基于亚马逊评论数据构建真实用户跨年多域行为记忆数据集
  • 在12个领域上测试14个主流大模型,发现现有记忆方法仍远未达用户期望
  • 适合研究长期记忆、个性化推荐与跨域建模的学者使用

大语言模型上下文窗口已扩展至百万词级别,但记忆能力评估仍局限于短会话合成对话。我们提出 extsc{MemoryCD},首个基于真实世界行为的用户中心、跨领域长时记忆基准,数据源自亚马逊评论数据集中多年持续的用户交互。不同于依赖脚本化人设生成合成数据的现有数据集, extsc{MemoryCD}追踪真实用户的长期跨域行为。我们构建了涵盖14个主流基础大模型与6种记忆方法基线的多维度长上下文记忆评估流水线,在4类个性化任务与12个不同领域上评估代理在单域与跨域场景下的真实用户行为模拟能力。分析表明,现有记忆方法在多个领域中距离用户满意度仍有显著差距,为跨域长期个性化提供了首个可量化评估的测试平台。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synthetic dialogues. We introduce \textsc{MemoryCD}, the first large-scale, user-centric, cross-domain memory benchmark derived from lifelong real-world behaviors in the Amazon Review dataset. Unlike existing memory datasets that rely on scripted personas to generate synthetic user data, \textsc{MemoryCD} tracks authentic user interactions across years and multiple domains. We construct a multi-faceted long-context memory evaluation pipeline of 14 state-of-the-art LLM base models with 6 memory method baselines on 4 distinct personalization tasks over 12 diverse domains to evaluate an agent's ability to simulate real user behaviors in both single and cross-domain settings. Our analysis reveals that existing memory methods are far from user satisfaction in various domains, offering the first testbed for cross-domain life-long personalization evaluation.

长时记忆个性化跨域建模评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。