arXiv:2507.14590cs.CLcs.AI2025-07被引 3

用大模型做情感分类数据增强,传统方法效果不输生成式。

Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification

  • 用GPT对文本进行反向翻译和改写来扩充数据。
  • 反向翻译和改写在多个任务中表现优于零样本/少样本生成。
  • 适合资源有限但需提升小样本情感分类性能的研究者。

众多特定领域机器学习任务面临数据稀缺和类别不平衡问题。本文系统研究了自然语言处理中的数据增强方法,尤其关注基于GPT等大语言模型的方法。旨在评估传统方法如改写与反向翻译,能否借助新一代模型实现与纯生成式方法相当甚至更优的性能。选取了针对数据稀缺问题并利用ChatGPT的方法,以及一个典型数据集,设计多组实验对比四种不同数据增强策略。通过生成数据质量及对分类性能的影响进行评估,结果显示:反向翻译与改写在多数情况下可达到或超越零样本与少样本生成的效果。

原文摘要 · Abstract (English)

Numerous domain-specific machine learning tasks struggle with data scarcity and class imbalance. This paper systematically explores data augmentation methods for NLP, particularly through large language models like GPT. The purpose of this paper is to examine and evaluate whether traditional methods such as paraphrasing and backtranslation can leverage a new generation of models to achieve comparable performance to purely generative methods. Methods aimed at solving the problem of data scarcity and utilizing ChatGPT were chosen, as well as an exemplary dataset. We conducted a series of experiments comparing four different approaches to data augmentation in multiple experimental setups. We then evaluated the results both in terms of the quality of generated data and its impact on classification performance. The key findings indicate that backtranslation and paraphrasing can yield comparable or even better results than zero and a few-shot generation of examples.

数据增强情感分类大模型NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。