arXiv:2608.14813cs.CL2026-08中稿 · the CPSS workshop …

发现大模型训练数据中含数十万极端言论文档,警示安全风险。

Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data

论文配图:Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
图 1 · 摘自论文原文
  • 用自动化+专家验证结合的方法检测训练数据
  • Dolma数据集至少含数十万极端内容文档
  • 适合关注AI安全与数据伦理的研究者

尽管学术界对可信和安全AI高度关注,但大语言模型在预训练和后训练阶段所接触的文本语料构成仍缺乏重视。本文针对大模型是否暴露于未经过滤、无上下文的极端言论问题展开研究。基于官方文件与学术文献中的多种极端言论定义,结合自动化文本处理与专家验证的提取流程,我们对支撑OLMo系列模型的开源训练语料Dolma进行了评估。结果表明,Dolma极有可能包含数十万份含有极端内容及仇恨言论的文档,涵盖直接暴力号召等类型。该发现揭示了数据构建与模型预训练中的潜在风险,呼吁加强数据筛选与治理。

原文摘要 · Abstract (English)

Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn much attention. In this work, we address the question of whether LLMs are exposed to unfiltered, uncontextualised extremist speech. Using several definitions of extremist speech, stemming from official documents and research literature, and an extraction pipeline combining automated text processing with expert verification, we provide a lower bound on the prevalence of extremist documents in Dolma, an open training corpus underpinning the OLMo series of models. We show that Dolma is likely to include hundreds of thousands of documents containing extremist content and hate speech of several types, including direct calls for violence, and discuss the implications of this for data curation and model pre-training.

大模型安全数据伦理极端言论AI风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。