用虚拟用户心理世界评估情感支持是否真正个性化
EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal World
- 构建三角色代理模拟用户内心,让系统自评支持效果
- 20个大模型在100个真实用户画像上测试,普遍缺乏个性化
- 适合研究情感计算、人机共情的开发者和学者使用
当前情感支持对话的评估范式倾向于奖励通用共情回应,却无法检验支持是否真正契合用户的独特心理特征与情境需求。我们提出 EmoHarbor,一个基于‘用户即评委’理念的自动化评估框架,通过模拟用户的内在心理世界实现评估。EmoHarbor 采用 Chain-of-Agent 架构,将用户内部过程分解为三个专业化角色,使代理能与支持者互动并完成评估,行为类似真人用户。我们基于100个真实用户画像构建该基准,涵盖多样人格特质与情境,并定义了10个个性化支持质量的评估维度。对20个先进大模型的全面评估发现:尽管这些模型擅长生成共情回应,但始终未能根据个体用户情境进行有效适配。这一发现重新定义了核心挑战,推动研究从提升通用共情转向发展真正具备用户感知能力的情感支持系统。EmoHarbor 提供了一个可复现、可扩展的框架,助力更精细、用户导向的情感支持系统研发。
原文摘要 · Abstract (English)
Current evaluation paradigms for emotional support conversations tend to reward generic empathetic responses, yet they fail to assess whether the support is genuinely personalized to users' unique psychological profiles and contextual needs. We introduce EmoHarbor, an automated evaluation framework that adopts a User-as-a-Judge paradigm by simulating the user's inner world. EmoHarbor employs a Chain-of-Agent architecture that decomposes users' internal processes into three specialized roles, enabling agents to interact with supporters and complete assessments in a manner similar to human users. We instantiate this benchmark using 100 real-world user profiles that cover a diverse range of personality traits and situations, and define 10 evaluation dimensions of personalized support quality. Comprehensive evaluation of 20 advanced LLMs on EmoHarbor reveals a critical insight: while these models excel at generating empathetic responses, they consistently fail to tailor support to individual user contexts. This finding reframes the central challenge, shifting research focus from merely enhancing generic empathy to developing truly user-aware emotional support. EmoHarbor provides a reproducible and scalable framework to guide the development and evaluation of more nuanced and user-aware emotional support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。