让分割模型更懂跨域少样本任务,性能提升超11%。
TAVP: Task-Adaptive Visual Prompt for Cross-domain Few-shot Segmentation
- 用多级特征融合+自适应提示生成,让模型灵活适配新任务。
- 在1次和5次示例下,准确率分别提升1.3%和11.76%。
- 适合做跨领域少样本图像分割的研究者与工程师。
尽管大视觉模型(LVM)在图像理解中展现出巨大潜力,但基于大规模预训练的段落任意模型(SAM)在跨域和少样本分割任务中表现仍不理想。现有方法难以将基础模型的知识有效迁移至新场景。为此,本文提出一种任务自适应视觉提示框架(TAVP),用于解决跨域少样本分割(CD-FSS)问题。首先,采用多级特征融合(MFF)提取综合先验特征;其次,引入类别-域无关的自适应提示模块(CDTAP),实现对类别与域信息的解耦,并生成高质量可学习的视觉提示。该方法通过生成式提示机制与专用原型计算,在保留SAM原有知识的同时,显著提升模型适应能力。在四个跨域数据集上的实验证明,本模型优于当前最优方法,在1次示例设置下平均准确率提升1.3%,5次示例下提升11.76%。
原文摘要 · Abstract (English)
While large visual models (LVM) demonstrated significant potential in image understanding, due to the application of large-scale pre-training, the Segment Anything Model (SAM) has also achieved great success in the field of image segmentation, supporting flexible interactive cues and strong learning capabilities. However, SAM's performance often falls short in cross-domain and few-shot applications. Previous work has performed poorly in transferring prior knowledge from base models to new applications. To tackle this issue, we propose a task-adaptive auto-visual prompt framework, a new paradigm for Cross-dominan Few-shot segmentation (CD-FSS). First, a Multi-level Feature Fusion (MFF) was used for integrated feature extraction as prior knowledge. Besides, we incorporate a Class Domain Task-Adaptive Auto-Prompt (CDTAP) module to enable class-domain agnostic feature extraction and generate high-quality, learnable visual prompts. This significant advancement uses a unique generative approach to prompts alongside a comprehensive model structure and specialized prototype computation. While ensuring that the prior knowledge of SAM is not discarded, the new branch disentangles category and domain information through prototypes, guiding it in adapting the CD-FSS. Comprehensive experiments across four cross-domain datasets demonstrate that our model outperforms the state-of-the-art CD-FSS approach, achieving an average accuracy improvement of 1.3\% in the 1-shot setting and 11.76\% in the 5-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。