arXiv:2602.20193cs.CRcs.AI2026-02ACL

攻击扩散模型编码器会引发无触发的语义漂移,破坏生成内容的本质结构。

When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks

  • 通过雅可比分析发现后门是低秩目标变形,放大局部敏感性导致语义传播失真。
  • 提出SEMAD框架,量化嵌入漂移与下游功能错位,揭示深层结构风险。
  • 适合关注生成模型安全、几何审计与对抗性污染的研究者阅读。

文本到图像(T2I)模型的后门攻击传统评估主要关注触发激活与视觉保真度。本文挑战这一范式,证明编码器侧污染会引发持续存在的、无需触发的语义腐败,从根本上重塑表示流形。我们通过雅可比分析揭示,后门表现为低秩、以目标为中心的形变,增强局部敏感性,导致失真在语义邻域间协同传播。为严格量化这种结构性退化,我们提出SEMAD(语义对齐与漂移)诊断框架,测量内部嵌入漂移与下游功能错位。研究结果在扩散与对比学习范式中均得到验证,暴露了编码器污染的深层结构风险,强调了超越简单攻击成功率的几何审计必要性。

原文摘要 · Abstract (English)

Standard evaluations of backdoor attacks on text-to-image (T2I) models primarily measure trigger activation and visual fidelity. We challenge this paradigm, demonstrating that encoder-side poisoning induces persistent, trigger-free semantic corruption that fundamentally reshapes the representation manifold. We trace this vulnerability to a geometric mechanism: a Jacobian-based analysis reveals that backdoors act as low-rank, target-centered deformations that amplify local sensitivity, causing distortion to propagate coherently across semantic neighborhoods. To rigorously quantify this structural degradation, we introduce SEMAD (Semantic Alignment and Drift), a diagnostic framework that measures both internal embedding drift and downstream functional misalignment. Our findings, validated across diffusion and contrastive paradigms, expose the deep structural risks of encoder poisoning and highlight the necessity of geometric audits beyond simple attack success rates.

扩散模型后门攻击语义漂移几何审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。