提出跨检查框架,统一检测各类学习模式下的后门攻击。
Lie Detector: Unified Backdoor Detection via Cross-Examination Framework
- 通过双服务商模型对比发现不一致,定位后门触发器。
- 在监督、半监督、自回归任务上分别提升5.4%、1.6%、11.9%准确率。
- 首次有效检测多模态大语言模型中的后门,适合安全训练场景。
资源有限的机构常将模型训练外包给第三方,在半诚实环境下假设其遵守预设学习范式(如监督或半监督学习)。然而,此做法可能引入严重安全风险:攻击者可通过污染训练数据,在模型中植入后门。现有检测方法多依赖统计分析,难以在不同学习范式下保持普适准确性。为此,我们提出一种统一的后门检测框架,利用两个独立服务提供商之间的模型不一致性进行交叉检查。具体地,融合中心核对齐技术,实现跨模型架构与学习范式的鲁棒特征相似性度量,从而精确恢复和识别后门触发器。进一步引入后门微调敏感性分析,有效区分后门触发器与对抗扰动,显著降低误报率。大量实验表明,该方法在监督、半监督及自回归学习任务上,相较最先进基线分别提升5.4%、1.6%、11.9%的检测准确率。尤为关键的是,它是首个在多模态大语言模型中有效检测后门的方法,凸显其广泛适用性,推动安全深度学习发展。
原文摘要 · Abstract (English)
Institutions with limited data and computing resources often outsource model training to third-party providers in a semi-honest setting, assuming adherence to prescribed training protocols with pre-defined learning paradigm (e.g., supervised or semi-supervised learning). However, this practice can introduce severe security risks, as adversaries may poison the training data to embed backdoors into the resulting model. Existing detection approaches predominantly rely on statistical analyses, which often fail to maintain universally accurate detection accuracy across different learning paradigms. To address this challenge, we propose a unified backdoor detection framework in the semi-honest setting that exploits cross-examination of model inconsistencies between two independent service providers. Specifically, we integrate central kernel alignment to enable robust feature similarity measurements across different model architectures and learning paradigms, thereby facilitating precise recovery and identification of backdoor triggers. We further introduce backdoor fine-tuning sensitivity analysis to distinguish backdoor triggers from adversarial perturbations, substantially reducing false positives. Extensive experiments demonstrate that our method achieves superior detection performance, improving accuracy by 5.4%, 1.6%, and 11.9% over SoTA baselines across supervised, semi-supervised, and autoregressive learning tasks, respectively. Notably, it is the first to effectively detect backdoors in multimodal large language models, further highlighting its broad applicability and advancing secure deep learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。