激活提示让视觉提示更高效,显著缩小与微调的性能差距。
Visual prompting reimagined: The power of the Activation Prompts
- 将提示从输入层扩展到中间特征层,实现模型内部激活图的通用扰动。
- 在29个数据集上验证,激活提示在准确率和效率上全面优于传统输入提示。
- 揭示了不同模型对提示层的依赖偏好,为高效轻量微调提供新思路。
视觉提示(VP)作为一种流行方法,可将预训练视觉模型适配至下游任务,通过直接在输入数据中引入通用扰动来实现任务特定调整,而不修改模型参数。然而,当前VP与传统微调之间仍存在明显性能差距,亟需理论与实践层面的深入探索以推动输入级提示的发展。为此,本文提出广义概念——激活提示(AP),将输入级提示的范围拓展至模型中间层的激活图,允许通用扰动作用于深层特征。通过重审视觉提示问题并将其作为分析工具,我们揭示了输入级提示在性能与效率上的内在局限,并发现激活提示表现出依赖于模型架构的层偏好。研究发现,激活提示与卷积神经网络及视觉变换器中的归一化调优密切相关,但不同模型类型具有不同的最优提示层。通过跨29个数据集与多种模型架构的广泛实验,我们系统评估了激活提示的表现,结果表明其在准确率、时间、参数、内存和吞吐量等多方面均优于传统输入提示与参数高效微调基线。
原文摘要 · Abstract (English)
Visual prompting (VP) has emerged as a popular method to repurpose pretrained vision models for adaptation to downstream tasks. Unlike conventional model fine-tuning techniques, VP introduces a universal perturbation directly into the input data to facilitate task-specific fine-tuning rather than modifying model parameters. However, there exists a noticeable performance gap between VP and conventional fine-tuning methods, highlighting an unexplored realm in theory and practice to understand and advance the input-level VP to reduce its current performance gap. Towards this end, we introduce a generalized concept, termed activation prompt (AP), which extends the scope of the input-level VP by enabling universal perturbations to be applied to activation maps within the intermediate layers of the model. By using AP to revisit the problem of VP and employing it as an analytical tool, we demonstrate the intrinsic limitations of VP in both performance and efficiency, revealing why input-level prompting may lack effectiveness compared to AP, which exhibits a model-dependent layer preference. We show that AP is closely related to normalization tuning in convolutional neural networks and vision transformers, although each model type has distinct layer preferences for prompting. We also theoretically elucidate the rationale behind such a preference by analyzing global features across layers. Through extensive experiments across 29 datasets and various model architectures, we provide a comprehensive performance analysis of AP, comparing it with VP and parameter-efficient fine-tuning baselines. Our results demonstrate AP's superiority in both accuracy and efficiency, considering factors such as time, parameters, memory usage, and throughput.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。