arXiv:2510.06397cs.LGcs.AI2025-10

利用双曲嵌入的曲率设计新型后门攻击,突破传统检测机制

Geometry-Aware Backdoor Attacks: Leveraging Curvature in Hyperbolic Embeddings

  • 基于双曲空间曲率设计自适应触发器,利用边界区域的敏感性差异
  • 攻击成功率随接近边界而提升,传统检测效果显著下降
  • 揭示几何特异性漏洞,为防御设计提供理论依据

非欧几里得基础模型越来越多地将表示置于如双曲几何等弯曲空间中。我们发现,这种几何结构会产生边界驱动的不对称性,后门触发器可加以利用:在边界附近,微小的输入变化对标准输入空间检测器而言看似不明显,却会在模型表示空间中引发不成比例的大范围偏移。我们的分析形式化了这一现象,并揭示了防御方法的局限性:通过沿径向向内拉点来抑制此类触发器的方法,虽有效但会牺牲模型在该方向上的有用敏感性。基于这些洞察,我们提出一种简单的几何自适应触发器,并在多种任务和架构上进行评估。实验表明,攻击成功率随接近边界而上升,而传统检测器性能则相应减弱,与理论趋势一致。这些结果揭示了非欧模型中的几何特异性漏洞,并为防御设计及其局限性提供了基于分析的指导。

原文摘要 · Abstract (English)

Non-Euclidean foundation models increasingly place representations in curved spaces such as hyperbolic geometry. We show that this geometry creates a boundary-driven asymmetry that backdoor triggers can exploit. Near the boundary, small input changes appear subtle to standard input-space detectors but produce disproportionately large shifts in the model's representation space. Our analysis formalizes this effect and also reveals a limitation for defenses: methods that act by pulling points inward along the radius can suppress such triggers, but only by sacrificing useful model sensitivity in that same direction. Building on these insights, we propose a simple geometry-adaptive trigger and evaluate it across tasks and architectures. Empirically, attack success increases toward the boundary, whereas conventional detectors weaken, mirroring the theoretical trends. Together, these results surface a geometry-specific vulnerability in non-Euclidean models and offer analysis-backed guidance for designing and understanding the limits of defenses.

后门攻击双曲嵌入模型安全几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。