arXiv:2409.14740cs.CLcs.AI2024-09EMNLP被引 24

用少量种子数据生成逼真有害内容,提升检测模型鲁棒性。

ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information

  • 基于少量种子数据生成多样且真实的有害文本
  • 在多个数据集上显著提升检测模型性能,接近人工标注水平
  • 适合需要增强模型抗干扰能力的研究者和安全团队

在不同自然语言处理任务中,检测有害内容对在线环境至关重要,尤其在社交媒体影响日益增长的背景下。然而,以往研究存在两大问题:1)低资源场景下数据匮乏;2)有害内容定义与判断标准不一致,导致分类模型需具备对虚假特征和多样化表现的鲁棒性。为此,我们提出ToxiCraft,一种新型有害信息合成框架,仅需少量种子数据即可生成大量多样、高度逼真的合成有害文本。在多个数据集上的实验表明,该框架显著提升了检测模型的鲁棒性和适应性,其性能逼近甚至超过人工标注的黄金标准。

原文摘要 · Abstract (English)

In different NLP tasks, detecting harmful content is crucial for online environments, especially with the growing influence of social media. However, previous research has two main issues: 1) a lack of data in low-resource settings, and 2) inconsistent definitions and criteria for judging harmful content, requiring classification models to be robust to spurious features and diverse. We propose Toxicraft, a novel framework for synthesizing datasets of harmful information to address these weaknesses. With only a small amount of seed data, our framework can generate a wide variety of synthetic, yet remarkably realistic, examples of toxic information. Experimentation across various datasets showcases a notable enhancement in detection model robustness and adaptability, surpassing or close to the gold labels.

有害内容生成数据合成模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。