arXiv:2508.05474cs.AIcs.CL2025-08被引 2

用小模型生成6个新对话情感数据集,提升分类效果

Can Large Language Models Generate Effective Datasets for Emotion Recognition in Conversations?

  • 用轻量级大模型合成多样化对话情感数据
  • 在3个主流基准上训练模型,性能显著提升
  • 适合需要数据增强的对话情感研究者

对话中的情感识别(ERC)旨在识别互动中的情绪变化,是推动机器智能的重要一步。然而,当前ERC数据稀缺,现有数据集因来源高度偏倚及软标签固有的主观性而面临诸多挑战。尽管大语言模型(LLMs)在情感任务中表现优异,但其训练成本高,且在ERC数据生成中的应用仍有限。为此,我们采用一个小型、资源高效且通用的LLM,合成具有多样特性的ERC数据集,补充三个最广泛使用的ERC基准。我们生成了六个新数据集,其中两个针对每个基准进行优化。实验评估了这些数据集在(1)补充现有数据以支持ERC分类,以及(2)分析标签不平衡影响方面的效用。结果表明,基于生成数据集训练的分类器展现出强鲁棒性,并在现有基准上持续实现统计显著的性能提升。

原文摘要 · Abstract (English)

Emotion recognition in conversations (ERC) focuses on identifying emotion shifts within interactions, representing a significant step toward advancing machine intelligence. However, ERC data remains scarce, and existing datasets face numerous challenges due to their highly biased sources and the inherent subjectivity of soft labels. Even though Large Language Models (LLMs) have demonstrated their quality in many affective tasks, they are typically expensive to train, and their application to ERC tasks--particularly in data generation--remains limited. To address these challenges, we employ a small, resource-efficient, and general-purpose LLM to synthesize ERC datasets with diverse properties, supplementing the three most widely used ERC benchmarks. We generate six novel datasets, with two tailored to enhance each benchmark. We evaluate the utility of these datasets to (1) supplement existing datasets for ERC classification, and (2) analyze the effects of label imbalance in ERC. Our experimental results indicate that ERC classifier models trained on the generated datasets exhibit strong robustness and consistently achieve statistically significant performance improvements on existing ERC benchmarks.

情感识别数据生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。