用测试时提示训练,让视觉大模型在少标注下实现精准医学影像分割。
Tuning Vision Foundation Model via Test-Time Prompt-Guided Training for VFSS Segmentations
- 通过点提示引导测试时自监督训练,利用数据增强解决提示歧义。
- 在VFSS-5k数据集上12个解剖结构平均Dice达0.868。
- 适合医疗影像领域,标注成本高但需高精度分割的场景。
视觉基础模型在通用和专业图像的分割任务中表现出卓越的泛化能力,但仍与特定任务模型存在性能差距。通常需在下游数据集上微调以缩小差距,但获取完全标注数据既困难又昂贵。为此,我们提出一种新型测试时训练范式,无需完整标注即可提升基础模型性能。方法利用简单点提示引导测试时半自监督训练任务,模型通过多种数据增强来消除点提示的歧义性。该方法直接应对医学影像领域标注耗时耗力的问题。我们在新构建的视频吞咽造影数据集(VFSS-5k)上进行广泛实验,针对实例分割任务,仅用单一模型即在12个解剖结构上取得平均Dice系数0.868。
原文摘要 · Abstract (English)
Vision foundation models have demonstrated exceptional generalization capabilities in segmentation tasks for both generic and specialized images. However, a performance gap persists between foundation models and task-specific, specialized models. Fine-tuning foundation models on downstream datasets is often necessary to bridge this gap. Unfortunately, obtaining fully annotated ground truth for downstream datasets is both challenging and costly. To address this limitation, we propose a novel test-time training paradigm that enhances the performance of foundation models on downstream datasets without requiring full annotations. Specifically, our method employs simple point prompts to guide a test-time semi-self-supervised training task. The model learns by resolving the ambiguity of the point prompt through various augmentations. This approach directly tackles challenges in the medical imaging field, where acquiring annotations is both time-intensive and expensive. We conducted extensive experiments on our new Videofluoroscopy dataset (VFSS-5k) for the instance segmentation task, achieving an average Dice coefficient of 0.868 across 12 anatomies with a single model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。