Setoka benchmark评估个性化代理对用户多层次理解能力,发现现有系统在深层认知理解上表现不足。
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

- 基于心理学理论构建四层用户理解体系:语义记忆、情景记忆、行为模式、人格特质。
- 实测显示模型在情景记忆任务上表现下降,跨源信息整合时性能进一步恶化。
- 适合研究个性化智能体、长期用户建模及隐私保护数据生成的学者使用。
个性化代理在各类任务中日益普及,有效辅助需不仅检索对话历史中的显式事实,还需推断抽象个人特征。然而,现有记忆基准主要评估显性信息检索能力,难以衡量深层用户理解。本文提出Setoka,一个基于异构数据的层次化用户理解评估基准。其基于认知与人格心理学理论,定义四个理解层级:语义记忆、情景记忆、行为模式与人格特质。为实现真实且隐私保护的评估,设计基于心理测量学的流水线,大规模合成多样、连贯的异构用户数据与查询。利用Setoka评估3个语言模型结合5个记忆系统对10个合成用户的性能。结果表明:尽管现有系统在语义记忆检索上表现良好,但在情景记忆上性能下降;面对需整合分散异构信息的行为模式与人格特质理解任务,性能进一步恶化。这说明用户理解不能仅靠简单事实检索,亟需支持跨源整合与长期行为抽象的记忆机制。
原文摘要 · Abstract (English)
Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。