提升视觉语言模型在零样本分类中的抗攻击能力。
Hierarchically Robust Zero-shot Vision-language Models

- 基于层级嵌入与多级对齐,增强图像文本模态的鲁棒性。
- 在多个数据集上实现更强的对抗鲁棒性,尤其对父类攻击有效。
- 适合关注模型安全性和层次化语义理解的研究者。
视觉语言模型(VLMs)虽可进行零样本分类,但易受对抗攻击影响。现有鲁棒微调方法通常将固定文本嵌入与图像嵌入对齐,导致自然性能与鲁棒性双重下降。当模型面临针对父类(如哺乳动物)而非仅基类(如猫)的对抗攻击时,鲁棒性进一步减弱。为此,我们提出一种基于层级嵌入的新型对抗微调框架,通过多层级图像-文本模态对齐,提升对抗鲁棒性。额外机制确保视觉嵌入位于层级中合适深度,并建立了嵌入深度与最大可行边缘大小之间的理论关联。模型天然支持多种边缘大小,有助于提升对抗泛化能力。由于不同树结构可共享相同叶节点标签,我们还考虑跨多棵树对齐以增强语义多样性。在多个数据集上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) can perform zero-shot classification but are susceptible to adversarial attacks. While robust fine-tuning improves their robustness, existing approaches align fixed text embeddings with an image embedding, sacrificing natural performance and robustness. A robustness degradation also occurs when a model faces adversarial attacks targeting superclasses (parent classes, e.g., mammal) in addition to their base (leaf) classes (e.g., cat). Thus, to enhance adversarial robustness and leverage the inherent hierarchical properties of class space, we propose a novel adversarial fine-tuning framework based on hierarchical embeddings and several levels of adversarially robust alignment of image-text modalities. Additional mechanisms place visual embeddings at the desired depth of hierarchy, and we provide a theoretical connection between the depth of embedding in the hierarchy and the maximum viable margin size. Our model naturally realizes several margin sizes, boosting generalization of adversaries for robustification. As various trees with different parent labels can share the same leaf labels, we also consider aligning over multiple trees to boost semantic variety. Experiments across several datasets are performed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。