arXiv:2509.05318cs.CRcs.AI2025-09

通过扰动不一致性检测预训练模型中的后门样本,无需额外数据或模型。

Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models

  • 利用掩码填充生成扰动,通过日志概率曲率衡量扰动差异的一致性。
  • 在四种典型攻击和五类大模型攻击中,检测准确率优于现有零样本方法。
  • 仅需预训练模型即可检测,适用于训练前和训练后阶段,实用性强。

未经审查的第三方和网络数据使预训练模型易受后门攻击。检测后门样本对防止推理时激活或训练时注入至关重要。然而,现有方法常需访问被污染模型、额外干净样本或大量计算资源,实用性受限。为此,我们提出基于扰动不一致性评估( ete)的后门样本检测方法。该方法可应用于预训练和后训练阶段。检测过程仅需现成的预训练模型计算样本对数概率,以及基于掩码填充策略的自动化函数生成扰动。该方法基于一个有趣现象:后门样本的扰动差异变化小于干净样本。据此,通过曲率衡量不同扰动样本与输入样本间对数概率的差异,评估扰动不一致性的稳定性,从而判断输入样本是否为后门样本。在四种典型后门攻击和五类大语言模型后门攻击上的实验表明,该策略显著优于现有零样本黑盒检测方法。

原文摘要 · Abstract (English)

The use of unvetted third-party and internet data renders pre-trained models susceptible to backdoor attacks. Detecting backdoor samples is critical to prevent backdoor activation during inference or injection during training. However, existing detection methods often require the defender to have access to the poisoned models, extra clean samples, or significant computational resources to detect backdoor samples, limiting their practicality. To address this limitation, we propose a backdoor sample detection method based on perturbatio\textbf{N} discr\textbf{E}pancy consis\textbf{T}ency \textbf{E}valuation (\NETE). This is a novel detection method that can be used both pre-training and post-training phases. In the detection process, it only requires an off-the-shelf pre-trained model to compute the log probability of samples and an automated function based on a mask-filling strategy to generate perturbations. Our method is based on the interesting phenomenon that the change in perturbation discrepancy for backdoor samples is smaller than that for clean samples. Based on this phenomenon, we use curvature to measure the discrepancy in log probabilities between different perturbed samples and input samples, thereby evaluating the consistency of the perturbation discrepancy to determine whether the input sample is a backdoor sample. Experiments conducted on four typical backdoor attacks and five types of large language model backdoor attacks demonstrate that our detection strategy outperforms existing zero-shot black-box detection methods.

后门检测大模型安全语言模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。