arXiv:2505.15592cs.CV2025-05

用少量数据快速适配视觉提示模型,提升专业领域分割精度

VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation

  • 融合多种高效微调方法,构建轻量级适配框架
  • 仅用5张标注图,分割性能提升50%(mIoU)
  • 支持实时交互,适合快速部署到新场景

大规模预训练视觉骨干网络通过强大的特征提取能力,推动了计算机视觉的发展,使无需训练的视觉提示方法在语义分割中得以应用。然而,在视觉特征与训练分布差异较大的专业领域,这类模型表现欠佳。为此,我们提出VP Lab,一个面向语义分割的迭代式增强视觉提示实验室。核心是E-PEFT——一种专为特定领域设计的参数高效微调技术集合,可实现低参数、低数据消耗的模型适配。该方法不仅超越了现有参数高效微调技术在Segment Anything Model(SAM)上的表现,还支持交互式、近实时的优化流程,用户可观察结果逐步提升。结合视觉提示与E-PEFT,在多个技术数据集上仅使用5张验证图像,即实现语义分割mIoU性能提升50%,建立了一种快速、高效、可交互的新范式。本工作以演示形式呈现。

原文摘要 · Abstract (English)

Large-scale pretrained vision backbones have transformed computer vision by providing powerful feature extractors that enable various downstream tasks, including training-free approaches like visual prompting for semantic segmentation. Despite their success in generic scenarios, these models often fall short when applied to specialized technical domains where the visual features differ significantly from their training distribution. To bridge this gap, we introduce VP Lab, a comprehensive iterative framework that enhances visual prompting for robust segmentation model development. At the core of VP Lab lies E-PEFT, a novel ensemble of parameter-efficient fine-tuning techniques specifically designed to adapt our visual prompting pipeline to specific domains in a manner that is both parameter- and data-efficient. Our approach not only surpasses the state-of-the-art in parameter-efficient fine-tuning for the Segment Anything Model (SAM), but also facilitates an interactive, near-real-time loop, allowing users to observe progressively improving results as they experiment within the framework. By integrating E-PEFT with visual prompting, we demonstrate a remarkable 50\% increase in semantic segmentation mIoU performance across various technical datasets using only 5 validated images, establishing a new paradigm for fast, efficient, and interactive model deployment in new, challenging domains. This work comes in the form of a demonstration.

视觉提示参数高效微调语义分割快速部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。