arXiv:2607.06643cs.CRcs.LG2026-07

通过吸收机制实现高效后门防御,零性能损失。

The Power of Backdoor Absorption in Community Training

论文配图:The Power of Backdoor Absorption in Community Training
图 1 · 摘自论文原文
  • 构建马尔可夫链模型分析恶意更新的吸收过程。
  • 仅10%步骤验证即可抑制后门,成功率趋近于零。
  • 适合资源受限的分布式训练场景使用。

后门攻击严重威胁大规模AI模型安全。在去中心化训练中,模型所有者将训练外包给外部计算提供者时,攻击者可设计隐蔽的低频触发器注入恶意行为,逃避常规审计。传统检测需重新完整计算训练过程,成本过高,违背所有者资源约束。本文研究连续优化动态对拜占庭扰动的鲁棒性,当攻击者控制f/n个训练者时,量化了模型所有者所需最小审计开销以概率性限制攻击成功。将注入-吸收动态形式化为离散时间马尔可夫链(DTMC),证明结合自然吸收、随机调度器与懒惰验证预言机的防御策略下,任何有界攻击者成功率渐近趋于零。实验证明,即使仅在10%训练步骤中调用验证预言机,也能显著抑制后门,且无性能损失。该方法为关键安全场景提供了可证明安全且计算高效的防御方案。

原文摘要 · Abstract (English)

Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to external compute providers within a decentralized training paradigm, adversaries can craft stealthy, low-frequency triggers to inject malicious behavior while evading standard audits. Traditionally, detecting these attacks requires a full re-computation of the training steps--a prohibitive overhead that directly contradicts the owner's resource constraints. To address this, we investigate the resilience of continuous optimization dynamics under Byzantine perturbations, where adversaries are forced to compete against a continuous influx of honest updates. Under a threat model where an adversary compromises f out of n total trainers, we quantify the minimum auditing overhead required by the model owner to probabilistically bound the attack success rate. We formalize this injection-absorption dynamic as a Discrete-Time Markov Chain (DTMC). Using this framework, we prove that the success probability of any bounded adversary asymptotically collapses to zero under a defense strategy combining natural absorption, a randomized scheduler, and lazy verification oracle. Empirical results demonstrate significant backdoor suppression with zero utility degradation even when invoking the verification oracle on merely 10% of the total training steps. This approach yields a provably sound and computationally efficient defense for safety-critical AI.

后门防御分布式训练马尔可夫链安全可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。