构建首个面向任务型对话的个性化评估基准,揭示大模型真实表现差异。
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
- 设计用户代理与裁判代理,模拟真实任务对话并评分。
- 多轮实验显示当前大模型个性化能力差异显著。
- 适合研究对话系统个性化、评估方法论的开发者和学者。
大型语言模型(LLMs)推动了对话式AI助手的发展,但系统性评估其个性化能力——即在完成任务时适应个体用户偏好——仍具挑战。现有个性化基准多聚焦于闲聊、非对话任务或特定领域,无法涵盖任务型个性化服务的复杂性。为此,我们提出PersonaLens,一个全面评估任务型AI助手个性化能力的基准。该基准包含多样化的用户档案,涵盖丰富偏好与交互历史,并配备两个基于LLM的专用代理:用户代理用于生成真实任务导向对话,裁判代理采用LLM-as-a-Judge范式评估个性化程度、回复质量与任务成功率。通过在多种任务上对当前主流LLM助手进行大规模实验,我们发现其个性化能力存在显著差异,为推进对话式AI系统提供了关键洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) have advanced conversational AI assistants. However, systematically evaluating how well these assistants apply personalization--adapting to individual user preferences while completing tasks--remains challenging. Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture the complexities of personalized task-oriented assistance. To address this, we introduce PersonaLens, a comprehensive benchmark for evaluating personalization in task-oriented AI assistants. Our benchmark features diverse user profiles equipped with rich preferences and interaction histories, along with two specialized LLM-based agents: a user agent that engages in realistic task-oriented dialogues with AI assistants, and a judge agent that employs the LLM-as-a-Judge paradigm to assess personalization, response quality, and task success. Through extensive experiments with current LLM assistants across diverse tasks, we reveal significant variability in their personalization capabilities, providing crucial insights for advancing conversational AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。