arXiv:2601.07023cs.AI2026-01ACL被引 10

用日记邮件等长期数字痕迹评估AI克隆的记忆能力

CloneMem: Benchmarking Long-Term Memory for AI Clones

  • 基于非对话类数字痕迹构建纵向记忆数据集
  • 现有记忆机制在追踪个人状态变化时表现不佳
  • 适合研究个性化AI与长期记忆的学者使用

AI克隆旨在模拟个体的思想与行为,实现长期个性化交互,对记忆系统提出了严苛要求,需建模随时间演进的经验、情感与观点。现有记忆基准多依赖碎片化的对话历史,难以捕捉连续的人生轨迹。我们提出CloneMem,一个基于非对话数字痕迹(如日记、社交媒体帖子、邮件)的长时记忆评估基准,覆盖1至3年跨度。该基准采用分层数据构建框架以确保时间连贯性,并定义任务评估智能体追踪个人状态演变的能力。实验表明当前记忆机制在此场景下表现不佳,揭示了面向生活化个性化AI的开放挑战。代码与数据集已开源。

原文摘要 · Abstract (English)

AI Clones aim to simulate an individual's thoughts and behaviors to enable long-term, personalized interaction, placing stringent demands on memory systems to model experiences, emotions, and opinions over time. Existing memory benchmarks primarily rely on user-agent conversational histories, which are temporally fragmented and insufficient for capturing continuous life trajectories. We introduce CloneMem, a benchmark for evaluating longterm memory in AI Clone scenarios grounded in non-conversational digital traces, including diaries, social media posts, and emails, spanning one to three years. CloneMem adopts a hierarchical data construction framework to ensure longitudinal coherence and defines tasks that assess an agent's ability to track evolving personal states. Experiments show that current memory mechanisms struggle in this setting, highlighting open challenges for life-grounded personalized AI. Code and dataset are available at https://github.com/AvatarMemory/CloneMemBench

AI克隆长期记忆数字痕迹评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。