用大模型生成混合语种数据,提升情感分析效果
Leveraging Large Language Models for Code-Mixed Data Augmentation in Sentiment Analysis
- 用大模型生成自然流畅的语码混杂文本
- 西英混合语中F1提升9.32%,马来语英混合需低基线时才有效
- 低成本生成真实感强的合成数据,适合资源少的语言
语码混杂(CM)在多语言社会中普遍存在,但因其复杂性和数据稀缺,给自然语言处理带来挑战。本文提出利用大语言模型生成合成的语码混杂数据,并用于提升特定任务模型在语码混杂情感分析中的性能。实验显示,在西班牙语-英语场景中,合成数据使F1分数提升9.32%,优于以往增强方法;而在马拉雅拉姆语-英语场景中,仅当基线表现较差时合成数据才有帮助;若已有较强自然数据,则额外合成数据收益有限。人工评估证实该方法能高效生成自然流畅的语码混杂句子,是一种简单且成本低廉的数据增强方式。研究结果表明,对大模型进行少样本提示是语码混杂数据增强的有前景方法,对提升情感分析性能具有重要意义,有助于社会影响力系统的建设。
原文摘要 · Abstract (English)
Code-mixing (CM), where speakers blend languages within a single expression, is prevalent in multilingual societies but poses challenges for natural language processing due to its complexity and limited data. We propose using a large language model to generate synthetic CM data, which is then used to enhance the performance of task-specific models for CM sentiment analysis. Our results show that in Spanish-English, synthetic data improved the F1 score by 9.32%, outperforming previous augmentation techniques. However, in Malayalam-English, synthetic data only helped when the baseline was low; with strong natural data, additional synthetic data offered little benefit. Human evaluation confirmed that this approach is a simple, cost-effective way to generate natural-sounding CM sentences, particularly beneficial for low baselines. Our findings suggest that few-shot prompting of large language models is a promising method for CM data augmentation and has significant impact on improving sentiment analysis, an important element in the development of social influence systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。