通过图像块增强与跨视图正则化,有效防御多模态大模型的后门攻击。
A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models
- 利用图像块级增强和跨视图输出差异正则化,识别并抑制后门响应。
- 在低中毒率下将攻击成功率降至不足5%,同时保持正常生成能力。
- 适合关注多模态模型安全、需应对隐蔽触发场景的研究者。
多模态大语言模型已成为统一处理视觉与语言任务的重要基础设施。然而,这类模型在监督微调过程中极易遭受后门植入,一旦触发特定模式,便会持续输出攻击者预设的有害响应。后门防御的核心挑战在于:在低中毒率下抑制攻击成功的同时,保持模型正常的生成能力,二者存在内在矛盾。强抑制常导致良性性能下降,弱正则化则无法有效缓解后门行为。为此,我们提出一种基于图像块增强与跨视图正则化的统一防御框架,从特征表示和输出分布两个层面协同约束模型对触发模式的异常响应。具体而言,结合图像块级数据增强与跨视图输出差异正则化,利用后门响应对非语义扰动异常不变的特性,主动拉大原始视图与扰动视图的输出分布差异,显著降低后门触发成功率。同时,通过输出熵约束避免过度抑制,保障正常指令生成质量。在三种模型、两项任务、六种攻击下的实验结果表明,该方法在有效降低攻击成功率的同时,维持了高水平的正常文本生成能力。本工作为大规模多模态模型在低频中毒与隐蔽触发场景下的安全可控部署提供了可行方案。
原文摘要 · Abstract (English)
Multimodal large language models have become an important infrastructure for unified processing of visual and linguistic tasks. However, such models are highly susceptible to backdoor implantation during supervised fine-tuning and will steadily output the attacker's predefined harmful responses once a specific trigger pattern is activated. The core challenge of backdoor defense lies in suppressing attack success under low poisoning ratios while preserving the model's normal generation ability. These two objectives are inherently conflicting. Strong suppression often degrades benign performance, whereas weak regularization fails to mitigate backdoor behaviors. To this end, we propose a unified defense framework based on patch augmentation and cross-view regularity, which simultaneously constrains the model's anomalous behaviors in response to triggered patterns from both the feature representation and output distribution levels. Specifically, patch-level data augmentation is combined with cross-view output difference regularization to exploit the fact that backdoor responses are abnormally invariant to non-semantic perturbations and to proactively pull apart the output distributions of the original and perturbed views, thereby significantly suppressing the success rate of backdoor triggering. At the same time, we avoid over-suppression of the model during defense by imposing output entropy constraints, ensuring the quality of normal command generation. Experimental results across three models, two tasks, and six attacks show that our proposed defense method effectively reduces the attack success rate while maintaining a high level of normal text generation capability. Our work enables the secure, controlled deployment of large-scale multimodal models in realistic low-frequency poisoning and covert triggering scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。