arXiv:2605.09549cs.LG2026-05

发现视觉语言模型自适应提示会失效,原因在于梯度失衡与门控退化。

When Adaptation Fails: A Gradient-Based Diagnosis of Collapsed Gating in Vision-Language Prompt Learning

论文配图:When Adaptation Fails: A Gradient-Based Diagnosis of Collapsed Gating in Vision-Language Prompt Learning
图 1 · 摘自论文原文
  • 通过梯度分析诊断自适应提示的失效机制。
  • 在多个数据集上验证门控模块输出趋同、梯度信号微弱。
  • 揭示参数高效学习中复杂结构未必有效,适合谨慎设计提示策略的人看。

自适应提示机制旨在通过动态调整提示以增强视觉语言模型性能。然而,在使用CLIP类主干网络的冻结少样本提示学习中,我们系统性观察到自适应门控和提示选择模块常发生崩溃:输出趋于恒定,梯度信号几乎消失,且性能往往不及固定提示。为深入探究此问题,我们在多个数据集和多种提示学习架构上开展控制实验,识别出两种反复出现的失败模式:梯度幅值失衡与门控退化。研究结果促使我们重新审视在参数高效学习中盲目增加架构复杂性的做法,并明确了在该范式下,提示级自适应门控何时有效、何时无效。

原文摘要 · Abstract (English)

Adaptive prompting mechanisms have been proposed to enhance vision-language models by dynamically tailoring prompts to inputs. However, in frozen few-shot prompt learning with CLIP-style backbones, we systematically observe that adaptive gates and prompt-selection modules often collapse: they produce nearly constant outputs, contribute negligible gradient signals, and frequently fail to outperform fixed prompts. To further explore this issue, we present a systematic diagnostic study to uncover the underlying causes and conditions of adaptation failure. Through controlled experiments across datasets and multiple prompt learning architectures, we identify two recurring failure modes: gradient magnitude imbalance and gate degradation. Our findings invite a re-examination of indiscriminately adding architectural complexity in parameter-efficient learning and clarify when prompt-level adaptive gating is, and is not, effective in this regime.

提示学习自适应视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。