用提示词增强生成冲突文本,提升敏感任务分类效果
PromptAug: Fine-grained Conflict Classification Using Data Augmentation
- 基于大模型提示工程生成冲突文本,规避内容安全限制
- 在极端数据稀缺下仍实现准确率与F1提升2%
- 结合定性分析发现四类生成缺陷,适合安全敏感领域研究者
随着社交媒体冲突增多,有效识别有害行为的分类模型至关重要。由于‘垃圾进垃圾出’原则,模型性能高度依赖训练数据质量。然而,针对冲突行为等细粒度任务的高质量标注数据稀缺、成本高且难以获取。加之平台日益限制研究数据访问,文本数据增强成为替代方案。但大模型防护机制阻碍了攻击性内容生成,使冲突数据增强面临独特挑战。本文提出PromptAug,一种创新的基于大模型的数据增强方法,在冲突与情绪数据集上分别实现准确率与F1-score提升2%。通过极端数据稀缺场景、定量多样性分析及定性主题分析,全面评估该方法。主题分析揭示四类生成问题:语言流畅性、幽默歧义、内容歧义和误解读。整体而言,PromptAug为冲突检测等敏感任务提供了有效的数据增强路径,融合自然语言处理与社会科学研究方法,具有跨学科价值。
原文摘要 · Abstract (English)
Given the rise of conflicts on social media, effective classification models to detect harmful behaviours are essential. Following the garbage-in-garbage-out maxim, machine learning performance depends heavily on training data quality. However, high-quality labelled data, especially for nuanced tasks like identifying conflict behaviours, is limited, expensive, and difficult to obtain. Additionally, as social media platforms increasingly restrict access to research data, text data augmentation is gaining attention as an alternative to generate training data. Augmenting conflict-related data poses unique challenges due to Large Language Model (LLM) guardrails that prevent generation of offensive content. This paper introduces PromptAug, an innovative LLM-based data augmentation method. PromptAug achieves statistically significant improvements of 2% in both accuracy and F1-score on conflict and emotion datasets. To thoroughly evaluate PromptAug against other data augmentation methods we conduct a robust evaluation using extreme data scarcity scenarios, quantitative diversity analysis and a qualitative thematic analysis. The thematic analysis identifies four problematic patterns in augmented text: Linguistic Fluidity, Humour Ambiguity, Augmented Content Ambiguity, and Augmented Content Misinterpretation. Overall, this work presents PromptAug as an effective method for augmenting data in sensitive tasks like conflict detection, offering a unique, interdisciplinary evaluation grounded in both natural language processing and social science methodology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。