提出动态评估框架,让个性化智能体适应用户长期需求变化。
Beyond Static Evaluation: Rethinking the Assessment of Personalized Agent Adaptability in Information Retrieval
- 用随时间演化的用户偏好模型模拟真实用户
- 通过上下文化访谈提取用户真实偏好
- 评估代理在多轮交互中行为改进能力,适合研究者与产品设计
个性化AI代理正成为现代信息检索的核心,但现有评估方法仍以静态基准和一次性指标为主,无法反映用户需求随时间演变的特性。这限制了我们对代理在长期、动态交互中是否真正适配个体的能力评估。本文提出一种重新思考自适应个性化评估的概念视角,将焦点从静态性能快照转向面向交互、持续演进的评估。该视角包含三个核心组件:(1) 基于人格建模的用户仿真,采用随时间演化的偏好模型;(2) 受参考访谈启发的结构化偏好获取协议,支持上下文中的偏好提取;(3) 能衡量代理在多轮会话与任务中行为改善的适应性评估机制。尽管近期研究已采用大模型驱动的用户仿真,本文将其置于更广泛的长期评估范式中。为验证观点,我们在电商搜索场景下使用PersonalWAB数据集开展案例研究。本文不仅呈现一个评估框架,更为理解与评价个性化作为持续、以用户为中心的过程奠定了概念基础。
原文摘要 · Abstract (English)
Personalized AI agents are becoming central to modern information retrieval, yet most evaluation methodologies remain static, relying on fixed benchmarks and one-off metrics that fail to reflect how users' needs evolve over time. These limitations hinder our ability to assess whether agents can meaningfully adapt to individuals across dynamic, longitudinal interactions. In this perspective paper, we propose a conceptual lens for rethinking evaluation in adaptive personalization, shifting the focus from static performance snapshots to interaction-aware, evolving assessments. We organize this lens around three core components: (1) persona-based user simulation with temporally evolving preference models; (2) structured elicitation protocols inspired by reference interviews to extract preferences in context; and (3) adaptation-aware evaluation mechanisms that measure how agent behavior improves across sessions and tasks. While recent works have embraced LLM-driven user simulation, we situate this practice within a broader paradigm for evaluating agents over time. To illustrate our ideas, we conduct a case study in e-commerce search using the PersonalWAB dataset. Beyond presenting a framework, our work lays a conceptual foundation for understanding and evaluating personalization as a continuous, user-centric endeavor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。