arXiv:2601.16880cs.LGcs.IT2026-01被引 1

揭示深度网络最小权重扰动机制,用于低秩触发后门攻击的可证明安全边界。

Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks

  • 推导单层权重扰动的精确公式,揭示输出变化所需的最小更新量。
  • 发现低秩压缩可稳定激活隐藏后门,且保持全精度模型准确率。
  • 提供可验证的参数更新下限,适用于后门攻击防御与分析。

本文推导了深度神经网络(DNNs)实现特定输出变化所需的最小范数权重扰动,并讨论决定其大小的因素。这些单层精确公式与基于多层利普希茨常数的通用鲁棒性保证进行对比,二者数量级相当,表明其在保障效果上具有相似性。该结果被应用于精度修改触发的后门攻击,建立了此类攻击无法成功的可证明压缩阈值;实验证明,低秩压缩可在不损失全精度准确率的前提下可靠激活隐含后门。这些表达式揭示了反向传播边缘如何控制逐层敏感性,并为与期望输出偏移一致的最小参数更新提供了可验证保证。

原文摘要 · Abstract (English)

The minimal norm weight perturbations of DNNs required to achieve a specified change in output are derived and the factors determining its size are discussed. These single-layer exact formulae are contrasted with more generic multi-layer Lipschitz constant based robustness guarantees; both are observed to be of the same order which indicates similar efficacy in their guarantees. These results are applied to precision-modification-activated backdoor attacks, establishing provable compression thresholds below which such attacks cannot succeed, and show empirically that low-rank compression can reliably activate latent backdoors while preserving full-precision accuracy. These expressions reveal how back-propagated margins govern layer-wise sensitivity and provide certifiable guarantees on the smallest parameter updates consistent with a desired output shift.

后门攻击权重扰动低秩压缩可验证安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。