arXiv:2605.28228cs.CL2026-05被引 1

测试聊天机器人在困难求助场景下的表现,发现多数系统难以应对真实情绪困境。

When Seekers Are Hard to Help: Evaluating Emotional Support Dialogue Systems in Worst-Case Interactions

论文配图:When Seekers Are Hard to Help: Evaluating Emotional Support Dialogue Systems in Worst-Case Interactions
图 1 · 摘自论文原文
  • 用专家模拟难搞的求助者,构建极端对话情境。
  • 17个系统在极端情况下表现普遍下降,最强模型也难维持互动与情绪改善。
  • 提出新评估框架和训练数据,帮助模型更抗压。

情感支持对话系统(ESDS)常通过大语言模型模拟求助者进行评估与训练,但这些模拟者通常合作、表达清晰、回应积极,导致评估结果过于乐观。本文研究极端困难情境下的系统表现,即求助者因低参与度、抗拒、自我披露少、情绪波动大或固执负面解读而难以被帮助。我们邀请八位资深心理咨询师模拟此类困难角色,与现有中文ESDS交互,提供评分并参与访谈。基于此,提炼出极端行为模式,并揭示当前系统的关键短板。随后提出一个基于大模型的极端情境模拟器及四项新评估指标:深层情绪理解、引导探索、情绪支持平衡性、真实且有依据的支持。对17个系统的评估显示,几乎所有模型在极端条件下性能显著下降。通用大模型整体更鲁棒,但即使最强模型也难以持续吸引用户并改善其情绪状态。最后,我们证明该模拟机制可生成有效训练数据,提升小型模型的韧性。

原文摘要 · Abstract (English)

Emotional Support Dialogue Systems (ESDSes) are increasingly evaluated and trained with LLM-simulated seekers. However, such simulated seekers often behave as cooperative, average-case users who disclose clearly, respond constructively, and accept support within a few turns. This can lead to overly optimistic evaluation and obscure whether ESDSes can handle difficult help-seeking interactions. In this work, we study ESDS evaluation under worst-case interactions, where seekers are hard to help due to low engagement, resistance, limited self-disclosure, emotional volatility, or rigid negative interpretations. We first conduct an expert simulation study with eight experienced counselling professionals, who simulate difficult seekers, interact with existing Chinese ESDSes, provide scale ratings, and participate in semi-structured interviews. Based on this study, we derive worst-case seeker behaviours and identify key limitations of current systems. We then propose a worst-case evaluation framework consisting of an LLM-based worst-case seeker simulator and four worst-case-oriented metrics: Deep Emotional Understanding, Guided Exploration, Balanced Emotional Support, and Authentic and Grounded Support. Evaluating 17 systems, we find that nearly all models suffer substantial performance drops under worst-case interactions. Large general-purpose LLMs are generally more robust than specialised ESDSes, but even the strongest models struggle to sustain engagement and improve seekers' emotional states. Finally, we show that worst-case simulation can also generate useful training data, improving the robustness of smaller models.

情感支持评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。