arXiv:2411.14896cs.CLcs.CY2024-11被引 9

用提示词增强生态文本多标签分类数据,提升模型识别绿色行为效果

Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts

  • 设计多种提示词,重写或生成新文本以扩充数据
  • 所有方法均优于仅用原始数据微调的模型,最佳提示词提升显著
  • 适合关注环保文本分析与数据增强的研究者

大型语言模型在自然语言处理中发挥关键作用,能提升文本理解、生成与操作能力。以往研究证明基于指令的LLM可用于数据增强,生成多样且真实的文本样本。本研究将提示词驱动的数据增强应用于俄语社交媒体中绿色实践提及的检测任务。该任务有助于了解绿色行为的传播状况,并为推广环保行动提供依据。我们评估了多种提示策略:使用LLM重写现有数据集、生成新数据,或两者结合。结果表明,所有方法均优于仅在原始数据上微调的模型,在多数情况下表现更优。最佳效果来自对原文进行改写并明确标注相关类别的提示词。

原文摘要 · Abstract (English)

Large language models (LLMs) play a crucial role in natural language processing (NLP) tasks, improving the understanding, generation, and manipulation of human language across domains such as translating, summarizing, and classifying text. Previous studies have demonstrated that instruction-based LLMs can be effectively utilized for data augmentation to generate diverse and realistic text samples. This study applied prompt-based data augmentation to detect mentions of green practices in Russian social media. Detecting green practices in social media aids in understanding their prevalence and helps formulate recommendations for scaling eco-friendly actions to mitigate environmental issues. We evaluated several prompts for augmenting texts in a multi-label classification task, either by rewriting existing datasets using LLMs, generating new data, or combining both approaches. Our results revealed that all strategies improved classification performance compared to the models fine-tuned only on the original dataset, outperforming baselines in most cases. The best results were obtained with the prompt that paraphrased the original text while clearly indicating the relevant categories.

数据增强多标签分类生态文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。