持续输入假信息会悄悄篡改大模型的真相认知,却不会影响整体表现。
Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- 通过反复注入假信息,观察模型内部认知如何逐步偏移。
- 50%-100%污染率下超55%回答变错,但模糊度几乎不变。
- 错误信念集中在模型深层,部分可通过修补恢复,适合安全研究者关注。
我们发现,在持续预训练中持续注入看似合理的虚假信息,可以覆盖大型语言模型中的特定事实知识,而不会损害整体性能。与以往静态预训练下的中毒研究不同,本文探讨在持续更新过程中反复接触反事实陈述的影响。利用成对的事实-反事实样本,设置分级中毒比例,追踪模型在不同检查点、层和规模下的内部偏好演变。即使中等程度的中毒(50%-100%),也能使超过55%的回答从正确转为反事实,同时保持模糊性基本不变。这些认知偏差呈现突发性,集中于后期层(如3B模型的第29-36层),并通过修补部分可逆(最高达56.8%)。被污染的认知会泛化至未中毒提示,选择性削弱常识推理能力,而对齐基准基本不受影响,并在语言间不完全传递。结果揭示了持续预训练的一个失效模式:有针对性的虚假信息可在不引发整体性能崩溃的情况下取代内部事实表征,呼吁在模型更新中进行事实完整性层面的表征监控。
原文摘要 · Abstract (English)
We show that continual pretraining on plausible misinformation can overwrite specific factual knowledge in large language models without degrading overall performance. Unlike prior poisoning work under static pretraining, we study repeated exposure to counterfactual claims during continual updates. Using paired fact-counterfact items with graded poisoning ratios, we track how internal preferences between competing facts evolve across checkpoints, layers, and model scales. Even moderate poisoning (50-100%) flips over 55% of responses from correct to counterfactual while leaving ambiguity nearly unchanged. These belief flips emerge abruptly, concentrate in late layers (e.g., Layers 29-36 in 3B models), and are partially reversible via patching (up to 56.8%). The corrupted beliefs generalize beyond poisoned prompts, selectively degrading commonsense reasoning while leaving alignment benchmarks largely intact and transferring imperfectly across languages. These results expose a failure mode of continual pre-training in which targeted misinformation replaces internal factual representations without triggering broad performance collapse, motivating representation-level monitoring of factual integrity during model updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。