用人工标注数据微调大模型,提升聊天机器人共情对话能力。
Are You Listening to Me? Fine-Tuning Chatbots for Empathetic Dialogue
- 基于专家手工标注的小规模共情对话数据集进行微调。
- 生成对话在情感结构上符合预期,但人类评估显示共情感不足。
- 强调需结合自动分析与人工评估,才能打造真正有温度的对话系统。
对话代理自ELIZA以来取得了显著进展,已广泛应用于医疗、教育和客户服务等领域。随着其深入融入日常人际互动,情感智能尤其是共情倾听能力变得愈发重要。本研究探讨了大语言模型(LLMs)在生成情感丰富对话时的表现。从专家手工构建的小规模共情对话数据集出发,利用ChatGPT和Gemini扩展对话内容。通过VADER情感分析与专家评估相结合的方式,分析对话中的情感演变。结果显示,生成对话虽在情感结构上与预期一致,但人类评估发现其共情感知度与连贯性存在明显差距。这表明,对话中的情绪建模不仅需要情感表达的结构匹配,还需具备质的深度,凸显了在开发情感智能代理时,自动化与人工评估方法相结合的重要性。
原文摘要 · Abstract (English)
Conversational agents have made significant progress since ELIZA, expanding their role across various domains, including healthcare, education, and customer service. As these agents become increasingly integrated into daily human interactions, the need for emotional intelligence, particularly empathetic listening, becomes increasingly essential. In this study, we explore how Large Language Models (LLMs) respond when tasked with generating emotionally rich interactions. Starting from a small dataset manually crafted by an expert to reflect empathic behavior, we extended the conversations using two LLMs: ChatGPT and Gemini. We analyzed the emotional progression of the dialogues using both sentiment analysis (via VADER) and expert assessments. While the generated conversations often mirrored the intended emotional structure, human evaluation revealed important differences in the perceived empathy and coherence of the responses. These findings suggest that emotion modeling in dialogues requires not only structural alignment in the expressed emotions but also qualitative depth, highlighting the importance of combining automated and humancentered methods in the development of emotionally competent agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。