arXiv:2510.13219cs.CV2025-10中稿 · TMLR 2026综述被引 29

系统梳理视觉提示适配技术,厘清轻量微调新范式

Prompt-based Adaptation in Large-scale Vision Models: A Survey

  • 按注入粒度分为像素级提示与标记级提示,按生成方式分固定/可学习/生成提示
  • 覆盖医学影像、三维点云、视觉语言等多场景应用,支持测试时自适应与可信AI
  • 首个全面综述视觉提示适配的论文,适合研究者快速掌握该领域脉络

在计算机视觉中,视觉提示(VP)和视觉提示调优(VPT)近年来成为大规模视觉模型在“预训练-微调”范式下全量微调的轻量且高效替代方案。然而,尽管进展迅速,其概念边界仍模糊不清,当前研究常混用VP与VPT,反映出对两类技术及其应用场景缺乏系统区分。本文从基础原理重新审视VP与VPT的设计,并将其统一于称为“提示适配”(Prompt-based Adaptation, PA)的框架中。在此框架下,我们根据注入粒度将方法区分为:像素级(VP)与标记级(VPT);根据生成机制分为固定、可学习与生成提示。此外,我们考察了PA在医学成像、3D点云、视觉语言任务等领域的集成应用,以及其在测试时自适应与可信人工智能中的作用。同时,总结了现有基准数据集,并指出现有挑战与未来方向。据我们所知,这是首个聚焦于提示适配方法及其特性的综合性综述,旨在为各领域研究人员与实践者提供清晰的研究路线图。

原文摘要 · Abstract (English)

In computer vision, Visual Prompting (VP) and Visual Prompt Tuning (VPT) have recently emerged as lightweight and effective alternatives to full fine-tuning for adapting large-scale vision models within the "pretrain-then-finetune" paradigm. However, despite rapid progress, their conceptual boundaries remain blurred, as VP and VPT are frequently used interchangeably in current research, reflecting a lack of systematic distinction between these techniques and their respective applications. In this survey, we revisit the designs of VP and VPT from first principles and conceptualize them within a unified framework termed Prompt-based Adaptation (PA). Within this framework, we distinguish methods based on their injection granularity: VP operates at the pixel level, while VPT injects prompts at the token level. We further categorize these methods by their generation mechanism into fixed, learnable, and generated prompts. Beyond the core methodologies, we examine PA integrations across diverse domains, including medical imaging, 3D point clouds, and vision-language tasks, as well as its role in test-time adaptation and trustworthy AI. We also summarize current benchmarks and identify key challenges and future directions. To the best of our knowledge, we are the first comprehensive survey dedicated to PA methodologies and applications in light of their distinct characteristics. Our survey aims to provide a clear roadmap for researchers and practitioners in all areas to understand and explore the evolving landscape of PA-related research.

视觉提示轻量微调综述模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。