arXiv:2502.17857cs.CL2025-02被引 7

用大模型自动生成10万条共情对话,无需人工标注。

SYNTHEMPATHY: A Scalable Empathy Corpus Generated Using LLMs Without Any Crowdsourcing

  • 用微调后的Mistral 7B生成共情对话数据。
  • 构建了包含10.5万条回复的SYNTHEMPATHY语料库。
  • 生成数据可提升模型共情能力,适合对话系统研究者。

已有研究表明,具备共情行为的语言模型更易被人类接受。尽管共情对构建有用对话代理至关重要,但可用于微调大模型的大规模共情对话语料库仍十分稀缺。现有语料库多依赖众包生成共情对话,成本高、耗时长且难以扩展。本文提出一种数据生成框架,构建了名为SYNTHEMPATHY的大规模共情语料库,包含10.5万条针对真实情境的共情回复,全部通过大模型生成。基于该语料库微调的Mistral 7B模型在平均共情评分上实现提升。

原文摘要 · Abstract (English)

Previous research has shown that humans are more receptive towards language models that that exhibit empathetic behavior. While empathy is essential for developing helpful dialogue agents, very few large corpora containing empathetic dialogues are available for fine-tune LLMs. The few existing corpora have largely relied on crowdsourcing to simulate empathetic conversations, a process that is expensive, time-consuming, and not scalable to larger datasets. We propose a data generation framework for developing SYNTHEMPATHY, a large corpus containing 105k empathetic responses to real-life situations compiled through LLM generation. A base Mistral 7B model fine-tuned on our SYNTHEMPATHY corpus exhibits an increase in the average empathy score.

共情生成大模型数据合成对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。