arXiv:2607.26099cs.CRcs.LG2026-07

研究后门攻击在训练与推理触发器不一致时的泛化能力,提出新框架提升攻击成功率。

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

论文配图:Lilith: Backdoor Generalization under Training-Inference Trigger Shift
图 1 · 摘自论文原文
  • 用单一训练锚点构建目标漏洞,再生成推理时可用的触发器家族。
  • 在多种数据集和模型上实现高成功率攻击,且对正常性能影响小。
  • 揭示了表示对齐是攻击成功关键,适合安全评估与防御研究者参考。

机器学习服务越来越多依赖公开数据、第三方提供商和外包训练,这为数据投毒攻击创造了机会,可在保持良性功能的同时植入持久恶意行为。然而,现有后门研究大多评估精确触发器复用、训练暴露的触发器多样性或预定义变换轴上的变化,因此留下一个关键盲区:从训练阶段触发器学到的后门能否泛化到训练中未出现的推理阶段触发器族。我们将其定义为训练-推理触发器偏移下的后门泛化问题,并提出Lilith——一种黑盒锚点到家族框架。仅使用分离的替代资源,Lilith首先通过单个训练锚点诱导出紧凑的目标侧漏洞,再构建一个仅限推理使用的、保持锚点诱导表示几何结构的有界触发器家族。我们通过锚点清除和家族可达性分析,推导出在局部正则性和有界替代-受害者差异下,家族级目标保留的充分条件。在多个数据集、架构、投毒率和防御机制下的实验表明,Lilith在有限性能损失下实现了高家族级攻击成功率,且触发器泛化差距小。额外分析显示,家族激活依赖于表示对齐而非生成机制,暴露了精确触发器评估所忽视的更广泛威胁。

原文摘要 · Abstract (English)

Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.

后门攻击安全评估泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。