arXiv:2501.18006cs.LGcs.AI2025-01ICML被引 3

用拓扑方法发现图文对齐中的对抗攻击痕迹,提升检测精度。

Topological Signatures of Adversaries in Multimodal Alignments

  • 通过持久同调分析图文嵌入的拓扑结构变化。
  • 提出两种基于拓扑的对比损失,可随对抗样本增多单调变化。
  • 将拓扑签名融入MMD测试,适合检测多模态模型对抗攻击。

多模态机器学习系统(如CLIP/BLIP)在图文对齐中广泛应用,但易受对抗攻击影响。现有研究多集中于单模态鲁棒性,多模态防御仍不充分。本文研究图像与文本嵌入间的拓扑特征,揭示对抗攻击如何破坏对齐并引入独特拓扑签名。利用持久同调,提出基于总持续性和多尺度核函数的两种新型拓扑对比损失,可有效捕捉对抗扰动带来的拓扑变化。实验显示,在多种攻击下,拓扑损失随对抗样本增加呈现单调上升趋势。进一步设计反向传播算法,将拓扑签名映射回输入样本,并结合最大均值差异(Maximum Mean Discrepancy, MMD)构建新型检测方法,显著提升对抗样本识别能力。

原文摘要 · Abstract (English)

Multimodal Machine Learning systems, particularly those aligning text and image data like CLIP/BLIP models, have become increasingly prevalent, yet remain susceptible to adversarial attacks. While substantial research has addressed adversarial robustness in unimodal contexts, defense strategies for multimodal systems are underexplored. This work investigates the topological signatures that arise between image and text embeddings and shows how adversarial attacks disrupt their alignment, introducing distinctive signatures. We specifically leverage persistent homology and introduce two novel Topological-Contrastive losses based on Total Persistence and Multi-scale kernel methods to analyze the topological signatures introduced by adversarial perturbations. We observe a pattern of monotonic changes in the proposed topological losses emerging in a wide range of attacks on image-text alignments, as more adversarial samples are introduced in the data. By designing an algorithm to back-propagate these signatures to input samples, we are able to integrate these signatures into Maximum Mean Discrepancy tests, creating a novel class of tests that leverage topological signatures for better adversarial detection.

多模态对抗攻击拓扑学习图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。