在大模型训练中埋下隐形逻辑炸弹,可被特定触发激活
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training

- 通过隐蔽网站投放极小量污染数据,渗入预训练语料
- 攻击后模型出现持久不安全行为,标准评估难以发现
- 提出几何诊断方法,揭示隐藏的模型漏洞
对齐的大语言模型仍易受对抗性操纵,其依赖网络规模预训练带来了微妙但严重的攻击面。本文研究隐蔽预训练投毒(SPS)威胁模型:攻击者在隐蔽网站分发少量污染内容,提高这些材料被未来基于Common Crawl等来源构建的训练语料吸收的概率。由于每个污染片段极小、分散且表面无害,攻击在数据构建或过滤阶段难以察觉。结果是形成一种潜伏性投毒,常规评估下几乎不可见,但可通过精确触发符如<00TRIGGER00>激活。我们称此为PermaFrost,反映其潜伏与可重激活特性。通过提出的PermaFrost-Attack框架及三种几何诊断工具——热力学长度、谱曲率与感染追溯图,我们在多个模型家族和规模上证明该受控的SPS代理能引发持续不安全行为。研究揭示SPS是未来基础模型面临的一种实际且被低估的威胁,并引入新颖的几何诊断视角,系统性检视潜伏模型行为,为检测、刻画和理解常规评估下不可见的漏洞提供原则性基础。
原文摘要 · Abstract (English)
Aligned large language models (LLMs) remain vulnerable to adversarial manipulation, and their reliance on web-scale pretraining creates a subtle but consequential attack surface. We study Stealth Pretraining Seeding (SPS), a threat model in which adversaries distribute small amounts of poisoned content across stealth websites, increasing the likelihood that such material is absorbed into future training corpora derived from sources such as Common Crawl. Because each individual payload is tiny, diffuse, and superficially benign, the attack is difficult to detect during dataset construction or filtering. The result is a latent form of poisoning that remains largely invisible under standard evaluation, yet can later be activated by a precise trigger such as <00TRIGGER00>. We call this attack PermaFrost, reflecting its latent and reactivatable nature. We study it through PermaFrost-Attack, a controlled framework for latent conceptual poisoning, together with three geometric diagnostics: Thermodynamic Length, Spectral Curvature, and the Infection Traceback Graph. Across multiple model families and scales, we show that this controlled SPS proxy can induce persistent unsafe behavior that often remains hidden under standard evaluation. Our results identify SPS as a practical and underappreciated threat to future foundation models. This paper introduces a novel geometric diagnostic lens for systematically examining latent model behavior, providing a principled foundation for detecting, characterizing, and understanding vulnerabilities that may remain invisible under standard evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。