通过双向动态反馈提升视觉语言模型抗干扰能力
Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models

- 设计闭环双向提示机制,利用文本与图像相互校正
- 在11个数据集上实现顶尖鲁棒性,且计算开销小
- 适合需要高可靠性、实时响应的多模态应用
视觉语言模型在下游任务中表现良好,但极易受到破坏跨模态语义对齐的对抗扰动。现有防御方法多为单向或静态结构,未能利用跨模态互补性及实例自适应保护。为此,本文提出闭环双向提示机制,将鲁棒性适配视为通过冻结编码器上的动态反馈环恢复跨模态一致性。引入语义锚作为稳定先验,约束循环更新,缓解扰动引起的特征失真。通过锚点引导的自举过程,文本语义可去噪视觉表征,而优化后的视觉信息则支持实例自适应提示更新,最终形成修正且鲁棒的一致共识。在11个数据集上的广泛评估验证了其最先进的鲁棒性及强泛化能力,同时在计算成本与准确率间保持良好权衡。
原文摘要 · Abstract (English)
Vision Language Models adapt well to downstream tasks but are highly vulnerable to adversarial perturbations that disrupt cross-modal semantic alignment. Existing defenses are largely unidirectional or structural, failing to exploit bidirectional cross-modal complementarity and instance-wise adaptive protection. To overcome the limitations of unidirectional and static defenses in adversarial settings, we propose Closed-Loop Bidirectional Prompting, casting robust adaptation as cross-modal agreement recovery via a dynamic feedback loop on frozen encoders. A Semantic Anchor is introduced as a stable prior to constrain cyclic updates and mitigate perturbation-induced feature corruption. Through anchor-based bootstrapping, textual semantics denoise visual representations, while the refined visuals enable instance-adaptive prompt updating, yielding a rectified and robust consensus. Extensive evaluations across 11 datasets validate state-of-the-art robustness and strong base-to-new generalization, while maintaining a favorable trade-off between computational cost and accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。