arXiv:2510.13992cs.SEcs.LG2025-10

检验代码模型后门检测中谱签名方法的有效性与改进空间

Signature in Code Backdoor Detection, how far are we?

  • 系统评估谱签名方法在代码模型后门检测中的表现
  • 发现现有设置常不最优,提出更优参数配置方案
  • 提出新代理指标,无需重训即可预估检测效果

随着大型语言模型(LLMs)日益融入软件开发流程,其也面临越来越多的对抗攻击。其中,后门攻击尤为严重,攻击者通过在训练数据中嵌入隐蔽触发器来操控模型输出。检测此类后门仍具挑战性,一种有前景的方法是利用谱签名防御,通过分析特征表示的特征向量识别被污染的数据。尽管已有研究探索了谱签名在神经网络中的应用,但近期研究表明其在代码模型上的效果可能不佳。本文重新审视谱签名方法在代码模型后门攻击中的适用性,系统评估其在多种攻击场景与防御配置下的有效性,分析其优劣。我们发现,当前广泛采用的谱签名设置通常并非最优。因此,我们进一步探究关键因素的不同设置影响,并发现一个新代理指标,可在无需模型重训练的情况下更准确地估计谱签名的实际性能。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become increasingly integrated into software development workflows, they also become prime targets for adversarial attacks. Among these, backdoor attacks are a significant threat, allowing attackers to manipulate model outputs through hidden triggers embedded in training data. Detecting such backdoors remains a challenge, and one promising approach is the use of Spectral Signature defense methods that identify poisoned data by analyzing feature representations through eigenvectors. While some prior works have explored Spectral Signatures for backdoor detection in neural networks, recent studies suggest that these methods may not be optimally effective for code models. In this paper, we revisit the applicability of Spectral Signature-based defenses in the context of backdoor attacks on code models. We systematically evaluate their effectiveness under various attack scenarios and defense configurations, analyzing their strengths and limitations. We found that the widely used setting of Spectral Signature in code backdoor detection is often suboptimal. Hence, we explored the impact of different settings of the key factors. We discovered a new proxy metric that can more accurately estimate the actual performance of Spectral Signature without model retraining after the defense.

后门检测代码模型谱签名LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。