arXiv:2505.02009cs.CLcs.LG2025-05IJCAI被引 11

分析并过滤网页数据中的有害内容,提升大模型预训练安全性。

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

  • 构建有害内容分类体系,区分主题性与毒性内容。
  • 开发高精度过滤模型HarmFormer和多危害评测基准HAVOC。
  • 适合关注AI伦理、模型安全与合规的开发者与研究者。

大型语言模型(LLMs)广泛应用于各类实际场景,其预训练依赖于海量网络数据集如Common Crawl、C4和FineWeb。这些数据虽提供高质量语言信息,却常含仇恨言论、虚假信息与偏见叙事等有害内容。在未经筛选的数据上训练模型,可能延续毒性行为、传播虚假信息并放大社会偏见,损害用户信任并引发伦理争议。本文对这些数据集中的不当内容进行大规模分析,提出涵盖意图的分类体系,将有害网页分为主题类与毒性类。我们构建了提示评估数据集TTP,设计高精度模型HarmFormer用于有害内容过滤,并提出多危害开放性毒性评测基准HAVOC,揭示模型对对抗性毒害输入的响应机制。我们公开了TTP、TTP-Eval、HAVOC及经HarmFormer推理后的C4样本。本工作为更安全的LLM预训练提供洞见,并支持负责任AI(RAI)合规实践。

原文摘要 · Abstract (English)

Large language models (LLMs) have become integral to various real-world applications, leveraging massive, web-sourced datasets like Common Crawl, C4, and FineWeb for pretraining. While these datasets provide linguistic data essential for high-quality natural language generation, they often contain harmful content, such as hate speech, misinformation, and biased narratives. Training LLMs on such unfiltered data risks perpetuating toxic behaviors, spreading misinformation, and amplifying societal biases which can undermine trust in LLM-driven applications and raise ethical concerns about their use. This paper presents a large-scale analysis of inappropriate content across these datasets, offering a comprehensive taxonomy that categorizes harmful webpages into Topical and Toxic based on their intent. We also introduce a prompt evaluation dataset, a high-accuracy Topical and Toxic Prompt (TTP), and a transformer-based model (HarmFormer) for harmful content filtering. Additionally, we create a new multi-harm open-ended toxicity benchmark (HAVOC) and provide crucial insights into how models respond to adversarial toxic inputs. We share TTP, TTP-Eval, HAVOC and a sample of C4 inferenced on HarmFormer. Our work offers insights into ensuring safer LLM pretraining and serves as a resource for Responsible AI (RAI) compliance.

模型安全有害内容LLM预训练RAI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。