用视觉大模型统一解决医疗图像无源无监督域适应问题
Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model
- 利用视觉大模型生成高质量伪标签,实现跨模态域自适应
- 在10个方向、22个解剖目标上达到当前最优性能
- 适合医疗影像领域需要快速适配新场景的研究者
无源无监督域适应(SFUDA)对深海学习模型在多样化临床场景中的部署至关重要。然而,现有方法通常针对低差异、特定域偏移设计,难以推广至统一的多模态、多目标框架,严重制约实际应用。为此,我们提出Tell2Adapt,一种利用视觉基础模型(VFM)泛化知识的新SFUDA框架。通过上下文感知提示正则化(CAPR),确保高保真提示生成,将多样文本提示转化为标准指令,从而高效生成高质量伪标签,指导轻量学生模型适配目标域。为保障临床可靠性,框架引入视觉合理性精炼(VPR),利用VFM的解剖学知识将模型预测重新锚定于目标图像的低层视觉特征,有效消除噪声与误报。我们在10个域适应方向和22个解剖目标(包括脑、心脏、息肉、腹部等)上开展迄今最全面的评估,结果表明Tell2Adapt在医疗图像分割任务中持续优于现有方法,成为统一型SFUDA的最新标杆。代码已开源。
原文摘要 · Abstract (English)
Source Free Unsupervised Domain Adaptation (SFUDA) is critical for deploying deep learning models across diverse clinical settings. However, existing methods are typically designed for low-gap, specific domain shifts and cannot generalize into a unified, multi-modalities, and multi-target framework, which presents a major barrier to real-world application. To overcome this issue, we introduce Tell2Adapt, a novel SFUDA framework that harnesses the vast, generalizable knowledge of the Vision Foundation Model (VFM). Our approach ensures high-fidelity VFM prompts through Context-Aware Prompts Regularization (CAPR), which robustly translates varied text prompts into canonical instructions. This enables the generation of high-quality pseudo-labels for efficiently adapting the lightweight student model to target domain. To guarantee clinical reliability, the framework incorporates Visual Plausibility Refinement (VPR), which leverages the VFM's anatomical knowledge to re-ground the adapted model's predictions in target image's low-level visual features, effectively removing noise and false positives. We conduct one of the most extensive SFUDA evaluations to date, validating our framework across 10 domain adaptation directions and 22 anatomical targets, including brain, cardiac, polyp, and abdominal targets. Our results demonstrate that Tell2Adapt consistently outperforms existing approaches, achieving SOTA for a unified SFUDA framework in medical image segmentation. Code are avaliable at https://github.com/derekshiii/Tell2Adapt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。