arXiv:2606.04660cs.CL2026-06

构建长期数字伴侣评估基准,测试跨会话记忆与情感适应能力。

LifeSide: Benchmarking Agents as Lifelong Digital Companions

论文配图:LifeSide: Benchmarking Agents as Lifelong Digital Companions
图 1 · 摘自论文原文
  • 以多轮记忆-情绪-环境循环为核心设计评估框架。
  • 2000个虚拟人格在11.1万任务中表现不佳,长期理解力严重不足。
  • 适合研究长期对话系统、数字伴侣的学者和开发者。

长期数字伴侣需整合跨会话线索,持续更新对用户认知,并适应隐私边界的动态变化。现有评估方法仅孤立测试记忆召回与短期共情能力,难以反映真实场景。为此,我们提出 enchmark,一个聚焦多会话 extit{Memory-Emotion-Environment} 循环的基准。通过将用户建模为具有分层资料与事件轨迹的持久世界,enchmark 利用多智能体仿真将环境动态映射至对话中,保留潜在想法与可观察表达之间的关键差距。在2000个虚拟人格和11.1万任务上,评估涵盖记忆追踪、用户理解、隐私控制与情感陪伴四项能力。实验结果揭示:即便模型在当前记忆基准上达到饱和,仍无法在长时程中维持准确的用户理解与真正的陪伴关系。

原文摘要 · Abstract (English)

Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries. Existing evaluations fail to capture this, testing memory recall and short-term empathy in isolation. To bridge this gap, we introduce \benchmark, a benchmark centered on multi-session \textit{Memory-Emotion-Environment} loops. By modeling users as persistent worlds with layered profiles and event trajectories, \benchmark uses multi-agent simulation to project environmental dynamics into dialogue, preserving the critical gap between latent thoughts and observable expressions. Evaluating 2,000 personas and 111K tasks across memory tracking, user understanding, privacy control, and emotional companionship, our experiment results reveal a stark reality: even models that saturate current memory benchmarks fail to sustain accurate user understanding and true companionship over long horizons.

数字伴侣多轮对话长期记忆评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。