评估大模型对话中持续契合人设的共情能力,突破传统单轮评价局限。
EMPA: Evaluating Persona-Aligned Empathy as a Process
- 将真实对话转化为可控心理场景,模拟长期共情干预过程
- 通过方向一致性、累积影响与稳定性三维度量化共情轨迹
- 适用于需长期响应、反馈稀疏的智能代理系统优化
基于大模型的对话代理在评估人设契合型共情时面临挑战:用户状态隐含难测,现场反馈稀疏且难以验证,看似支持的回应仍可能使对话轨迹偏离人设需求。我们提出EMPA,一种面向过程的评估框架,将共情视为持续干预而非孤立回复。EMPA将真实互动提炼为可控制的心理学基础情景,结合开放式多智能体沙盒,暴露策略适应与失败模式,并在潜在心理空间中通过方向对齐性、累积影响与稳定性评分轨迹。该框架生成可复现的信号与指标,支持长时程共情行为的对比与优化,亦适用于受隐式动态和弱反馈影响的其他智能体场景。
原文摘要 · Abstract (English)
Evaluating persona-aligned empathy in LLM-based dialogue agents remains challenging. User states are latent, feedback is sparse and difficult to verify in situ, and seemingly supportive turns can still accumulate into trajectories that drift from persona-specific needs. We introduce EMPA, a process-oriented framework that evaluates persona-aligned support as sustained intervention rather than isolated replies. EMPA distills real interactions into controllable, psychologically grounded scenarios, couples them with an open-ended multi-agent sandbox that exposes strategic adaptation and failure modes, and scores trajectories in a latent psychological space by directional alignment, cumulative impact, and stability. The resulting signals and metrics support reproducible comparison and optimization of long-horizon empathic behavior, and they extend to other agent settings shaped by latent dynamics and weak, hard-to-verify feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。