评测大模型在对话中主动影响用户的能力,强调个性化信息的重要性。
$Ψ$-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues

- 设计三种真实场景,用用户历史对话生成人格画像进行评测
- 主流模型说服力有限,顶尖模型仍有提升空间
- 提供用户画像可使性能平均提升18.24%,适合研究主动型AI
个性化是现代语言代理的关键能力。然而,现有研究多将个性化代理视为被动响应用户偏好的存在,限制了其主动提供建议或引导互动的能力。为系统评估此类主动个性化在真实交互中的表现,我们提出Ψ-Bench,一个评估大语言模型通过对话影响真实用户的基准。该基准设计了三个现实交互场景,结合对话历史生成显式用户画像,赋予模拟客户人格特征。我们在Ψ-Bench上评估了10个前沿大模型,发现尽管多数模型能生成连贯合理的论点,但即使是顶尖模型在说服力方面仍存在较大改进空间。此外,提供客户画像可带来平均18.24%的性能提升,凸显用户特定信息对有效说服的重要性。整体而言,本工作强调了人格敏感性影响作为评估和开发更主动个性化大模型的重要方向。代码已开源:https://github.com/Hanpx20/Psi-Bench。
原文摘要 · Abstract (English)
Personalization is a crucial capability of modern language agents. However, current research primarily positions personalized agents as passive responders to user preferences, limiting their ability to interact with users and provide suggestions or guidance proactively. To systematically evaluate such proactive personalization in realistic interactions, we propose $Ψ$-Bench, a benchmark for assessing LLMs' ability to influence realistic users through conversation. We design three real-world interaction scenarios that involve persuasion in $Ψ$-Bench, and endow simulated clients with personal characteristics through explicit user profiles derived from dialogue histories. We evaluate 10 frontier LLMs on $Ψ$-Bench and find that while most models can produce coherent and reasonable arguments, even state-of-the-art models still leave considerable room for improvement in persuasion. We also find that providing access to client profiles yields an average performance gain of 18.24\%, highlighting the importance of user-specific information for effective persuasion. Overall, our work highlights persona-sensitive influencing as a challenging yet practical direction for evaluating and developing more proactive personalized LLM agents. Codes are available at: https://github.com/Hanpx20/Psi-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。