让聊天机器人在对话中不重复使用相同回应策略,提升共情效果。
Discourse Diversity in Multi-Turn Empathic Dialogue

- 用强化学习优化多轮对话中的回应策略多样性
- 使大模型对话重复率降低26.3%,共情评分提升25.3%
- 适用于真实情感支持场景,解决模型套路化问题
大型语言模型在单轮对话中虽被评价为高度共情(Ayers et al., 2023;Lee et al., 2024),但其常表现出模式化倾向,重复使用相同的词汇、句法和话语结构(Jiang et al., 2025;Shaib et al., 2024;Namuduri et al., 2025)。现有研究较少关注这种模式是否延伸至话语动作层面——即回应对对方实际起到了什么作用。这一问题在共情对话中尤为重要,因有效支持需随对话推进采用多样化策略(Stiles et al., 1998)。已有研究表明,在单轮情境下,模型比人类更频繁重复相同策略序列(Gueorguieva et al., 2026)。本研究将分析扩展至多轮对话,发现模式化现象加剧:一旦某个策略出现在支持者回应中,模型在下一回应中重复该策略的概率接近人类的两倍(0.50–0.56 对比 0.27)。此现象在真实情感支持对话中普遍存在于各类大模型中,且无法被标准相似性度量捕捉。为此,我们提出MINT(多轮话语动作新颖性训练),首个面向多轮共情对话优化话语动作多样性的强化学习框架。最佳版本结合共情质量奖励与跨轮次策略新颖性信号,在1.7B和4B模型上使整体共情评分相比基线提升25.3%,同时在4B模型上将跨轮次话语动作重复率降低26.3%,优于所有基线方法(包括仅优化质量或词级多样性的方法)。结果表明,当前模型并非缺乏共情能力,而是缺乏在对话过程中灵活变换回应策略的能力。
原文摘要 · Abstract (English)
Large language models (LLMs) produce responses rated as highly empathic in single-turn settings (Ayers et al., 2023; Lee et al., 2024), yet they are also known to be formulaic generators that reuse the same lexical patterns, syntactic templates, and discourse structures across tasks (Jiang et al., 2025; Shaib et al., 2024; Namuduri et al., 2025). Less attention has been paid to whether this formulaicity extends to the level of discourse moves, i.e., what a response does for the person it is addressing. This question is especially consequential for empathic dialogue, where effective support demands not just a kind response at one moment but varied strategies as a conversation unfolds (Stiles et al., 1998). Indeed, prior work shows that LLMs reuse the same tactic sequences more than human supporters in single-turn settings (Gueorguieva et al., 2026). We extend this analysis to multi-turn conversations and find that the rigidity compounds: once a tactic appears in a supporter turn, LLMs reuse it in the next at nearly double the rate of humans (0.50-0.56 vs. 0.27). This pattern holds across LLMs serving as supporters in real emotional support conversations, and is invisible to standard similarity metrics. To address this gap, we introduce MINT (Multi-turn Inter-tactic Novelty Training), the first reinforcement learning framework to optimize discourse move diversity across multi-turn empathic dialogue. The best MINT variant combines an empathy quality reward with a cross-turn tactic novelty signal, improving aggregate empathy by 25.3% over vanilla across 1.7B and 4B models while reducing cross-turn discourse move repetition by 26.3% on the 4B model, surpassing all baselines including quality-only and token-level diversity methods on both measures. These results suggest that what current models lack is not empathy itself, but the ability to vary their discourse moves across a conversation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。