arXiv:2607.15267cs.AIcs.CL2026-07

利用公开讨论接口污染大规模预训练数据,可能植入难以察觉的恶意行为。

Pretraining Data Can Be Poisoned through Computational Propaganda

论文配图:Pretraining Data Can Be Poisoned through Computational Propaganda
图 1 · 摘自论文原文
  • 通过公开讨论界面注入恶意内容,实现对预训练数据的攻击。
  • 提出半衰期分析法(HalfLife),可估算爬虫数据中恶意内容的残留比例。
  • 揭示第三方网页是语言模型预训练阶段潜在的攻击入口,适合安全研究者关注。

污染预训练数据可能引入难以检测和缓解的有害行为。以往研究多针对维基百科等固定数据源,这些数据规模小且同质性强,无法代表真实预训练语料的广泛异质性,且忽略了中毒数据与数据清洗流程的交互影响。本文通过现有大规模网络内容注入机制——公开讨论接口,证明在更复杂环境下实施预训练数据污染攻击是可行的。为评估爬虫采集与数据清洗后恶意内容是否留存,我们提出一种新分析方法 HalfLife,用于估算基于网络爬取的语言模型训练数据中对抗性内容的渗入程度。实验表明,该方法能有效识别毒化风险,确立第三方网页内容作为语言模型预训练攻击的新路径。

原文摘要 · Abstract (English)

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.

数据污染语言模型安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。