让视觉语言模型在有噪声标签时仍能精准少样本学习
Noise-aware few-shot learning through bi-directional multi-view prompt alignment

- 通过双向多视角提示对齐,区分干净与噪声信号
- 在合成和真实噪声数据集上显著超越现有方法
- 适合需要抗噪声少样本学习的场景
视觉语言模型通过提示调优具备强大的少样本能力,但对噪声标签敏感,易导致提示污染和跨模态对齐退化。现有方法因难以建模细粒度语义线索且无法自适应分离干净与噪声信号而受限。为此,我们提出NA-MVP框架,实现噪声感知的少样本学习。其核心思想是:鲁棒提示学习需从全局匹配转向区域感知对齐,明确区分干净与噪声线索。NA-MVP采用三项机制:(1)结合非平衡最优传输的多视角提示,实现细粒度图像块到提示的对应关系,同时抑制不可靠区域;(2)双向提示设计,捕捉互补的清洁导向与噪声感知线索,使模型聚焦于稳定语义;(3)基于对齐的定向优化策略,利用最优传输仅修正误标样本,保留可靠数据。在合成与真实世界噪声基准测试中,NA-MVP持续优于先进基线,验证了其在噪声监督下实现鲁棒少样本学习的有效性。
原文摘要 · Abstract (English)
Vision-language models offer strong few-shot capability through prompt tuning but remain vulnerable to noisy labels, which can corrupt prompts and degrade cross-modal alignment. Existing approaches struggle because they often lack the ability to model fine-grained semantic cues and to adaptively separate clean from noisy signals. To address these challenges, we propose NA-MVP, a framework for Noise-Aware few-shot learning through bi-directional Multi-View Prompt alignment. NA-MVP is built upon a key conceptual shift: robust prompt learning requires moving from global matching to region-aware alignment that explicitly distinguishes clean cues from noisy ones. To realize this, NA-MVP employs (1) multi-view prompts combined with unbalanced optimal transport to achieve fine-grained patch-to-prompt correspondence while suppressing unreliable regions; (2) a bi-directional prompt design that captures complementary clean-oriented and noise-aware cues, enabling the model to focus on stable semantics; and (3) an alignment-guided selective refinement strategy that uses optimal transport to correct only mislabeled samples while retaining reliable data. Experiments on synthetic and real-world noisy benchmarks demonstrate that NA-MVP consistently outperforms state-of-the-art baselines, confirming its effectiveness in enabling robust few-shot learning under noisy supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。