用普通词汇作触发器,让后门攻击延迟激活,更难被发现。
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
- 通过非线性衰减机制延迟后门触发,实现可控延迟期
- 保持94%以上干净准确率,激活后攻击成功率接近99%
- 可绕过现有主流防御,适合研究模型安全的学者
针对预训练模型(PTM)的后门攻击传统上假设恶意行为在触发后立即显现。本文提出一种新型威胁——延迟后门攻击(DBA),其激活与触发暴露在时间上解耦。设计并实现原型DND,通过轻量级有状态逻辑模块,将攻击激活延迟至达到设定阈值后才爆发,形成明确的延迟阶段和可控的爆发期。建立形式化模型刻画延迟行为,并提出双指标评估框架(ASR与ASR$_{delay}$)量化延迟效果。在四个NLP基准测试中验证:攻击可在可控时间内保持休眠,维持≥94%的干净准确率,激活后攻击成功率≈99%(其他方法平均低于95%),且对多种先进防御具有鲁棒性。本研究首次实证表明时间维度是预训练模型中一个可行但未受保护的新攻击面,亟需下一代具备状态感知和时序意识的防御机制。
原文摘要 · Abstract (English)
Backdoor attacks against pre-trained models (PTMs) have traditionally operated under an ``immediacy assumption,'' where malicious behavior manifests instantly upon trigger occurrence. This work revisits and challenges this paradigm by introducing \textit{\textbf{Delayed Backdoor Attacks (DBA)}}, a new class of threats in which activation is temporally decoupled from trigger exposure. We propose that this \textbf{temporal dimension} is the key to unlocking a previously infeasible class of attacks: those that use common, everyday words as triggers. To examine the feasibility of this paradigm, we design and implement a proof-of-concept prototype, termed \underline{D}elayed Backdoor Attacks Based on \underline{N}onlinear \underline{D}ecay (DND). DND embeds a lightweight, stateful logic module that postpones activation until a configurable threshold is reached, producing a distinct latency phase followed by a controlled outbreak. We derive a formal model to characterize this latency behavior and propose a dual-metric evaluation framework (ASR and ASR$_{delay}$) to empirically measure the delay effect. Extensive experiments on four (natural language processing)NLP benchmarks validate the core capabilities of DND: it remains dormant for a controllable duration, sustains high clean accuracy ($\ge$94\%), and achieves near-perfect post-activation attack success rates ($\approx$99\%, The average of other methods is below 95\%.). Moreover, DND exhibits resilience against several state-of-the-art defenses. This study provides the first empirical evidence that the temporal dimension constitutes a viable yet unprotected attack surface in PTMs, underscoring the need for next-generation, stateful, and time-aware defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。