提出多标签评测框架,更精准检测大模型生成内容的多重毒性。
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
- 构建三个基于15类毒性的多标签评测集,覆盖真实场景复杂毒性。
- 实验显示新方法在多标签检测上超越GPT-4o和DeepSeek等先进模型。
- 证明伪标签训练优于单标签监督,降低标注成本并提升性能。
大语言模型在自然语言处理任务中表现卓越,但生成有害内容的风险引发严重安全担忧。现有毒性检测主要依赖单标签基准,难以捕捉现实毒性内容固有的模糊性与多维度特征,导致检测偏差,包括漏检与误报,削弱了检测器可靠性。此外,细粒度毒性类别下的多标签标注成本极高,制约评估与模型发展。为此,我们构建了三个新型多标签毒性检测基准:Q-A-MLL、R-A-MLL 和 H-X-MLL,源自公开数据集,并依据详细15类毒性分类体系进行标注。我们进一步提供理论证明,在所发布数据集上,使用伪标签训练比直接单标签监督表现更优。同时,我们提出一种基于伪标签的毒性检测方法。大量实验表明,该方法显著优于GPT-4o和DeepSeek等先进基线,实现对大模型生成内容多标签毒性的更准确、可靠评估。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved impressive results across a range of natural language processing tasks, but their potential to generate harmful content has raised serious safety concerns. Current toxicity detectors primarily rely on single-label benchmarks, which cannot adequately capture the inherently ambiguous and multi-dimensional nature of real-world toxic prompts. This limitation results in biased evaluations, including missed toxic detections and false positives, undermining the reliability of existing detectors. Additionally, gathering comprehensive multi-label annotations across fine-grained toxicity categories is prohibitively costly, further hindering effective evaluation and development. To tackle these issues, we introduce three novel multi-label benchmarks for toxicity detection: \textbf{Q-A-MLL}, \textbf{R-A-MLL}, and \textbf{H-X-MLL}, derived from public toxicity datasets and annotated according to a detailed 15-category taxonomy. We further provide a theoretical proof that, on our released datasets, training with pseudo-labels yields better performance than directly learning from single-label supervision. In addition, we develop a pseudo-label-based toxicity detection method. Extensive experimental results show that our approach significantly surpasses advanced baselines, including GPT-4o and DeepSeek, thus enabling more accurate and reliable evaluation of multi-label toxicity in LLM-generated content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。