arXiv:2511.09003cs.CL2025-11被引 6

用动态轨迹评估大模型情感支持能力,更真实反映长期陪伴效果。

Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language Models

  • 构建基于情绪轨迹的评估框架,关注长期情感变化
  • 引入3个新指标,量化情绪稳定性与改善程度
  • 适合研究人机情感交互、心理陪伴的开发者参考

情感支持是人机交互的核心能力,应用于心理咨询、角色扮演和陪伴等场景。现有大语言模型(LLMs)评估多依赖短时静态对话,难以捕捉情感支持的动态与长期特性。为此,本文从快照式评估转向轨迹式评估,采用用户中心视角,衡量模型在对话过程中改善并稳定用户情绪状态的能力。框架构建包含328个情感情境与1,152个扰动事件的大规模基准数据集,模拟真实情绪演变过程。通过引入经验证的情绪调节策略(如情境选择与认知重评)约束模型输出,并将用户情绪轨迹建模为一阶马尔可夫过程,结合因果调整的情绪估计方法实现无偏情绪追踪。在此基础上,提出三个轨迹级指标:基线情绪水平(BEL)、情绪轨迹波动性(ETV)和情绪中心位置(ECP),综合刻画用户情绪随时间的变化特征,支持对大模型长期情感支持能力的全面评估。跨多种大模型的实证分析揭示显著差异,为模型优化提供可操作洞见。

原文摘要 · Abstract (English)

Emotional support is a core capability in human-AI interaction, with applications including psychological counseling, role play, and companionship. However, existing evaluations of large language models (LLMs) often rely on short, static dialogues and fail to capture the dynamic and long-term nature of emotional support. To overcome this limitation, we shift from snapshot-based evaluation to trajectory-based assessment, adopting a user-centered perspective that evaluates models based on their ability to improve and stabilize user emotional states over time. Our framework constructs a large-scale benchmark consisting of 328 emotional contexts and 1,152 disturbance events, simulating realistic emotional shifts under evolving dialogue scenarios. To encourage psychologically grounded responses, we constrain model outputs using validated emotion regulation strategies such as situation selection and cognitive reappraisal. User emotional trajectories are modeled as a first-order Markov process, and we apply causally-adjusted emotion estimation to obtain unbiased emotional state tracking. Based on this framework, we introduce three trajectory-level metrics: Baseline Emotional Level (BEL), Emotional Trajectory Volatility (ETV), and Emotional Centroid Position (ECP). These metrics collectively capture user emotional dynamics over time and support comprehensive evaluation of long-term emotional support performance of LLMs. Extensive evaluations across a diverse set of LLMs reveal significant disparities in emotional support capabilities and provide actionable insights for model development.

情感支持评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。