arXiv:2506.02326cs.CLcs.AI2025-06ACL被引 3

构建大规模毒性语言标注数据集,助力更精准识别毒害内容与目标群体。

Something Just Like TRuST : Toxicity Recognition of Span and Target

  • 设计多阶段人工标注流程,确保30万条标注中1.1万条高质量数据。
  • 发现微调的预训练模型在三项任务上均优于大模型,推理模型未提升性能。
  • 适合研究安全语言模型、内容审核与社会感知生成技术的学者使用。

毒性语言包括冒犯性、侮辱性或煽动伤害的内容。防止大语言模型(LLMs)产生有害输出的进展受限于对毒性定义的不一致。本文提出TRuST,一个大规模数据集,通过精心设计的毒性定义和标注方案统一并扩展了现有资源,包含约30万条标注,其中约1.1万条经高质量人工标注。为保证质量,设计了严格的多阶段人工标注流程,并评估了标注者多样性。在三个任务上对当前主流的LLMs和预训练模型进行了基准测试:毒性检测、目标群体识别和有毒词汇定位。结果表明,微调的预训练模型在三项任务中表现均优于大模型,而当前推理模型未能可靠提升性能。TRuST是评估与缓解LLM毒性最全面的资源之一,可推动社会敏感型与更安全的语言技术研究。

原文摘要 · Abstract (English)

Toxic language includes content that is offensive, abusive, or that promotes harm. Progress in preventing toxic output from large language models (LLMs) is hampered by inconsistent definitions of toxicity. We introduce TRuST, a large-scale dataset that unifies and expands prior resources through a carefully synthesized definition of toxicity, and corresponding annotation scheme. It consists of ~300k annotations, with high-quality human annotation on ~11k. To ensure high-quality, we designed a rigorous, multi-stage human annotation process, and evaluated the diversity of the annotators. Then we benchmarked state-of-the-art LLMs and pre-trained models on three tasks: toxicity detection, identification of the target group, and of toxic words. Our results indicate that fine-tuned PLMs outperform LLMs on the three tasks, and that current reasoning models do not reliably improve performance. TRuST constitutes one of the most comprehensive resources for evaluating and mitigating LLM toxicity, and other research in socially-aware and safer language technologies.

毒性检测数据集安全生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。