arXiv:2412.16905cs.CRcs.AI2024-12被引 1

通过修改模型结构植入隐蔽后门,触发器肉眼难辨且检测失效。

Backdoor Attack with Invisible Triggers Based on Model Architecture Modification

  • 修改预训练模型架构嵌入后门,无需污染训练数据。
  • 触发器完全隐形,既无法人工发现也逃过主流检测工具。
  • 适用于对模型安全性要求高的场景,如自动驾驶与医疗诊断。

机器学习系统易受后门攻击,攻击者可通过数据污染或架构修改操纵模型行为。传统攻击依赖在训练数据中注入含特定触发器的恶意样本,使模型在遇到对应触发器时输出目标错误结果。更复杂的攻击直接修改模型架构,将后门嵌入其中,从而规避基于数据的检测方法。然而,现有架构类攻击仍需可见触发器才能激活后门。本文提出一种新方法,将后门嵌入模型架构内,并能生成无迹可寻、极难察觉的触发器。该攻击通过修改预训练模型并重新分发实现,对不知情用户构成潜在威胁。在标准计算机视觉基准上的全面实验验证了攻击的有效性,其触发器在人工视觉检查及先进检测工具下均未被识别。

原文摘要 · Abstract (English)

Machine learning systems are vulnerable to backdoor attacks, where attackers manipulate model behavior through data tampering or architectural modifications. Traditional backdoor attacks involve injecting malicious samples with specific triggers into the training data, causing the model to produce targeted incorrect outputs in the presence of the corresponding triggers. More sophisticated attacks modify the model's architecture directly, embedding backdoors that are harder to detect as they evade traditional data-based detection methods. However, the drawback of the architectural modification based backdoor attacks is that the trigger must be visible in order to activate the backdoor. To further strengthen the invisibility of the backdoor attacks, a novel backdoor attack method is presented in the paper. To be more specific, this method embeds the backdoor within the model's architecture and has the capability to generate inconspicuous and stealthy triggers. The attack is implemented by modifying pre-trained models, which are then redistributed, thereby posing a potential threat to unsuspecting users. Comprehensive experiments conducted on standard computer vision benchmarks validate the effectiveness of this attack and highlight the stealthiness of its triggers, which remain undetectable through both manual visual inspection and advanced detection tools.

后门攻击模型安全隐蔽触发器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。