用少量种子数据生成逼真有害内容,提升检测模型鲁棒性。
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
- 基于少量种子数据生成多样且真实的有害文本
- 在多个数据集上显著提升检测模型性能,接近人工标注水平
- 适合需要增强模型抗干扰能力的研究者和安全团队
在不同自然语言处理任务中,检测有害内容对在线环境至关重要,尤其在社交媒体影响日益增长的背景下。然而,以往研究存在两大问题:1)低资源场景下数据匮乏;2)有害内容定义与判断标准不一致,导致分类模型需具备对虚假特征和多样化表现的鲁棒性。为此,我们提出ToxiCraft,一种新型有害信息合成框架,仅需少量种子数据即可生成大量多样、高度逼真的合成有害文本。在多个数据集上的实验表明,该框架显著提升了检测模型的鲁棒性和适应性,其性能逼近甚至超过人工标注的黄金标准。
原文摘要 · Abstract (English)
In different NLP tasks, detecting harmful content is crucial for online environments, especially with the growing influence of social media. However, previous research has two main issues: 1) a lack of data in low-resource settings, and 2) inconsistent definitions and criteria for judging harmful content, requiring classification models to be robust to spurious features and diverse. We propose Toxicraft, a novel framework for synthesizing datasets of harmful information to address these weaknesses. With only a small amount of seed data, our framework can generate a wide variety of synthetic, yet remarkably realistic, examples of toxic information. Experimentation across various datasets showcases a notable enhancement in detection model robustness and adaptability, surpassing or close to the gold labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。