arXiv:2508.04677cs.CV2025-08

通过引入弱语义噪声提升视觉语言模型的鲁棒性与泛化能力

Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise

  • 主动引入弱语义噪声,通过聚类构建噪声提示
  • 在11个基准上实现优于现有方法的鲁棒性和泛化性能
  • 适合需要抗噪声、跨类别泛化的视觉语言任务

Prompt tuning虽表现优异,但对未见类别仍缺乏鲁棒性和泛化能力。实验表明,完全消除语义噪声是限制鲁棒性的关键因素。现有方法通常在提示空间中抑制或过滤语义噪声,反而损害模型鲁棒性与泛化能力。为此,我们提出ANPrompt框架,主动引入弱语义噪声。通过将弱扰动特征聚类为噪声提示,并与文本和视觉编码器中的可学习标记结合,实现对语义变化的可控暴露。为增强视觉路径,提出噪声抵抗视觉提示原型(NRVPP),在弱扰动下稳定视觉语义。此外,在logits层面设计弱对齐损失(WALoss),强制清洁与扰动预测间的一致性,提供稳定监督。通过弱语义噪声暴露与基于logits的一致性联合优化,ANPrompt防止过拟合特定表述,同时保持语义完整性。在11个基准(含base-to-new划分)上的大量实验表明,ANPrompt持续优于现有prompt tuning方法,展现出更强的语义噪声鲁棒性与任务泛化能力。

原文摘要 · Abstract (English)

Prompt tuning has shown promising results, but its robustness and generalization to unseen categories remain limited. Through our experiments, we demonstrate that the complete removal of semantic noise is a key factor restricting robustness. Existing methods typically suppress or filter out semantic noise in the prompt space, inadvertently hindering the model's robustness and its ability to generalize to unseen categories. To address this, we propose ANPrompt, a robust prompt tuning framework that actively incorporates weak semantic noise. By clustering weakly perturbed features into noise prompts and integrating them with learnable tokens in both the text and vision encoders, ANPrompt ensures controlled exposure to semantic variations. To enhance the visual pathway, we introduce the Noise-Resistant Visual Prompt Prototype (NRVPP), which stabilizes visual semantics under weak perturbations. Additionally, we propose a Weak Alignment Loss (WALoss) at the logits level to enforce consistency between clean and perturbed predictions, providing stable supervision. By combining weak semantic noise exposure with logits-based consistency, ANPrompt prevents overfitting to specific phrasings while preserving semantic integrity. Extensive experiments across 11 benchmarks, including base-to-new splits, show that ANPrompt consistently outperforms existing prompt tuning methods, offering superior robustness to semantic noise and improved generalization across tasks.

提示调优视觉语言鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。