arXiv:2608.21194cs.CV2026-08

用能量引导动态生成图像专属提示,高效适配模型且泛化更强。

ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation

论文配图:ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation
图 1 · 摘自论文原文
  • 基于低秩初始化与能量引导,动态生成每张图专属提示。
  • 在15个数据集上优于当前最佳方法,参数量仅为1/590。
  • 适合追求高效、通用模型适配的研究者和工程师。

视觉提示(Visual Prompting, VP)已成为高效适配预训练模型至下游任务的方法。然而,现有方法在灵活性与效率间存在权衡:部分方法对所有图像使用固定提示,忽略个体差异;另一些方法引入辅助网络生成多样化提示,虽提升性能却显著增加参数量并易过拟合。此外,辅助网络与预训练模型固有偏差限制了可扩展性与泛化能力。本文提出能量引导的动态视觉提示(ES-VP),通过低秩初始化与能量引导的动态适应机制,生成图像特定提示,在参数更少的情况下实现更优性能。ES-VP直接利用预训练模型进行提示自适应生成,兼顾参数效率与泛化能力。在五个架构、十五个数据集上的实验表明,其持续超越当前最优单提示与多样提示方法。例如,在四个数据集上使用CLIP架构时,相比SOTA方法DAM-VP,平均准确率提升2.6%,同时仅需590倍更少的提示参数,确立了高效且通用模型适配的新基准。

原文摘要 · Abstract (English)

Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches encounter a trade-off between flexibility and efficiency. Some methods apply a fixed prompt to all images, ignoring individual image characteristics, while others introduce auxiliary networks to generate diverse prompts. Although the latter can improve performance, it also significantly increases parameter usage and the potential for overfitting to specific datasets. Furthermore, the auxiliary networks, combined with inherent biases in pre-trained models, limit scalability and generalization. In this paper, we propose Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods. ES-VP directly utilizes the pre-trained model for adaptive prompt generation, ensuring both parameter efficiency and improved generalization. Extensive experiments conducted on five architectures across fifteen datasets demonstrate that ES-VP consistently outperforms current state-of-the-art (SOTA) single and diverse VP methods. For instance, using the CLIP architecture across four datasets, ES-VP outperforms the SOTA method DAM-VP by an average of 2.6\% in accuracy while utilizing 590$\times$ fewer VP parameters, thereby establishing a new benchmark for efficient and generalizable model adaptation.

视觉提示模型适配参数效率动态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。