用因果推理提升医学影像分割的跨域泛化能力
Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
- 结合CLIP与因果推断,识别并移除设备/成像差异等干扰因素
- 在多个医疗影像数据集上实现更高分割精度和更强跨域适应性
- 适合关注医学AI泛化性的研究者和临床应用开发者
视觉语言模型(如CLIP)在计算机视觉任务中展现出出色的零样本能力,但在医学影像领域仍面临挑战,主要因医学数据存在显著的域偏移,由设备差异、操作伪影和成像模式等混杂因素引起,导致模型在未见域上表现不佳。为此,本文提出多模态因果驱动表征学习(MCDRL)框架,将因果推断与视觉语言模型结合,以解决医学图像分割中的域泛化问题。MCDRL分两步实施:首先利用CLIP的跨模态能力通过文本提示识别候选病灶区域,并构建包含域特异性变异的混淆因子词典;其次训练一个因果干预网络,利用该词典识别并消除这些域特异性变异的影响,同时保留对分割任务至关重要的解剖结构信息。大量实验表明,MCDRL持续优于现有方法,在多个数据集上均取得更优的分割准确率,表现出强鲁棒性和泛化能力。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging remains challenging due to the high variability and complexity of medical data. Specifically, medical images often exhibit significant domain shifts caused by various confounders, including equipment differences, procedure artifacts, and imaging modes, which can lead to poor generalization when models are applied to unseen domains. To address this limitation, we propose Multimodal Causal-Driven Representation Learning (MCDRL), a novel framework that integrates causal inference with the VLM to tackle domain generalization in medical image segmentation. MCDRL is implemented in two steps: first, it leverages CLIP's cross-modal capabilities to identify candidate lesion regions and construct a confounder dictionary through text prompts, specifically designed to represent domain-specific variations; second, it trains a causal intervention network that utilizes this dictionary to identify and eliminate the influence of these domain-specific variations while preserving the anatomical structural information critical for segmentation tasks. Extensive experiments demonstrate that MCDRL consistently outperforms competing methods, yielding superior segmentation accuracy and exhibiting robust generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。