用大模型生成毒液文本训练净化模型,效果反而更差。
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
- 用 Llama 3 等模型生成毒语替代人工标注数据
- 合成数据训练的模型性能比人类数据低最多 30%
- 大模型生成毒语词汇重复,缺乏多样性
现代大语言模型(LLMs)在生成合成数据方面表现出色,但在敏感领域如文本去毒方面尚未得到足够关注。本文探讨使用 LLM 生成的合成毒语数据作为训练去毒模型的替代方案。基于 ParaDetox 与 SST-2 数据集中的中性文本,利用 Llama 3 与 Qwen 激活修补模型生成对应毒语。实验表明,使用合成数据微调的模型在联合指标上表现持续低于人类标注数据训练的模型,性能下降最高达 30%。根本原因在于词汇多样性严重不足:大模型生成的毒语仅使用少量重复性侮辱词汇,无法涵盖人类毒语的复杂性和多样性。研究揭示了当前大模型在该领域的局限性,强调多样且人工标注的数据对构建鲁棒去毒系统仍至关重要。
原文摘要 · Abstract (English)
Modern Large Language Models (LLMs) are excellent at generating synthetic data. However, their performance in sensitive domains such as text detoxification has not received proper attention from the scientific community. This paper explores the possibility of using LLM-generated synthetic toxic data as an alternative to human-generated data for training models for detoxification. Using Llama 3 and Qwen activation-patched models, we generated synthetic toxic counterparts for neutral texts from ParaDetox and SST-2 datasets. Our experiments show that models fine-tuned on synthetic data consistently perform worse than those trained on human data, with a drop in performance of up to 30% in joint metrics. The root cause is identified as a critical lexical diversity gap: LLMs generate toxic content using a small, repetitive vocabulary of insults that fails to capture the nuances and variety of human toxicity. These findings highlight the limitations of current LLMs in this domain and emphasize the continued importance of diverse, human-annotated data for building robust detoxification systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。