arXiv:2411.15175cs.CLcs.AI2024-11被引 6

用开源大模型生成有毒数据,提升内容审核训练效果

ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?

  • 用提示工程与微调技术控制生成毒性内容
  • 微调后模型在5个数据集上减少幻觉与重复
  • 适合想低成本构建审核数据的团队使用

有效的有害内容检测依赖高质量且多样的数据,这是构建稳健内容审核模型的基础。合成数据已成为NLP各类任务中训练模型的常见方法。然而,对于仇恨言论检测这类高度主观的任务,其有效性仍不确定,以往研究结果不一。本研究探索开源大模型在生成有害数据方面的潜力,采用受控提示和监督微调技术以提升数据质量和多样性。我们系统评估了6个开源LLM在5个数据集上的表现,考察其生成多样化、高质量有害内容的能力,同时最小化幻觉和重复。结果表明,Mistral在所有开源模型中表现最优,且监督微调显著提升了数据的可靠性与多样性。我们进一步分析了基于提示与微调的合成方式之间的权衡,讨论实际部署挑战,并强调伦理考量。研究发现,经微调的开源大模型可为有毒内容检测数据集提供可扩展、低成本的增强方案,推动更开放透明的内容审核工具发展。

原文摘要 · Abstract (English)

Effective toxic content detection relies heavily on high-quality and diverse data, which serve as the foundation for robust content moderation models. Synthetic data has become a common approach for training models across various NLP tasks. However, its effectiveness remains uncertain for highly subjective tasks like hate speech detection, with previous research yielding mixed results. This study explores the potential of open-source LLMs for harmful data synthesis, utilizing controlled prompting and supervised fine-tuning techniques to enhance data quality and diversity. We systematically evaluated 6 open source LLMs on 5 datasets, assessing their ability to generate diverse, high-quality harmful data while minimizing hallucination and duplication. Our results show that Mistral consistently outperforms other open models, and supervised fine-tuning significantly enhances data reliability and diversity. We further analyze the trade-offs between prompt-based vs. fine-tuned toxic data synthesis, discuss real-world deployment challenges, and highlight ethical considerations. Our findings demonstrate that fine-tuned open source LLMs provide scalable and cost-effective solutions to augment toxic content detection datasets, paving the way for more accessible and transparent content moderation tools.

大模型生成内容审核合成数据开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。