arXiv:2608.05683cs.CVcs.AI2026-08

提出概率化多模态对齐方法,提升医学图像分割的不确定性感知能力。

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

论文配图:DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation
图 1 · 摘自论文原文
  • 通过概率适配器建模视觉与文本表征的不确定性
  • 在8个基准上仅用630万参数实现领先性能
  • 适合需要鲁棒性与泛化能力的临床医学图像分析

跨模态视觉与文本表征对齐是多模态医学图像理解的核心,但在真实临床条件下,两种模态均存在不确定性。现有视觉-语言分割方法依赖确定性跨模态匹配,忽视了边界模糊带来的偶然不确定性(aleatoric)和训练数据有限导致的认知不确定性(epistemic),导致域偏移下性能脆弱。为此,我们提出DistMedVL,一种基于冻结编码器的概率化视觉-语言框架,引入轻量级概率交叉模态适配器(PCM-Adapter),显式建模表示不确定性。PCM-Adapter包含两个序列模块:首先设计马哈拉诺比对齐模块(MAM),将文本标记建模为高斯分布,通过马哈拉诺比距离计算块-文本兼容性,实现方差条件匹配,降低不可靠特征维度权重;其次设计分布流模块(DFM),估计模态专属置信度参数,并进行视觉引导的文本分布优化,适应不同成像模态间的分布差异。在八个医学分割基准上的大量实验表明,DistMedVL仅使用630万可训练参数,即超越现有最优方法,在数据效率、扰动鲁棒性和跨数据集泛化方面表现优异。

原文摘要 · Abstract (English)

Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions. Existing vision-language segmentation methods rely on deterministic cross-modal matching, which overlooks aleatoric uncertainty from ambiguous boundaries and epistemic uncertainty from limited training data, leading to fragile performance under domain shift. To address this issue, we propose DistMedVL, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty. Specifically, the PCM-Adapter comprises two sequential modules for progressive probabilistic alignment. We first devise a Mahalanobis Alignment Module (MAM) that models textual tokens as Gaussian distributions and computes patch-text compatibility via Mahalanobis distance, yielding variance-conditioned matching that downweights unreliable feature dimensions. Moreover, we devise a Distribution Flow Module (DFM) that estimates modality-wise confidence parameters and performs vision-guided refinement of textual distributions, accommodating distributional variation across imaging modalities. Extensive experiments across eight medical segmentation benchmarks demonstrate that DistMedVL outperforms state-of-the-art methods with only 6.3M trainable parameters, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.

医学图像分割多模态对齐不确定性建模概率深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。