arXiv:2603.00156cs.CV2026-03被引 1

让医学影像分割更抗干扰,文字与图像双向互校正

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

  • 文字与图像双向迭代优化,提升语义对齐精度
  • 仅用1%标注数据仍保持高精度,抗运动模糊和低剂量噪声
  • 适合标注少、设备差的临床真实场景

医学图像分割是辅助诊断与治疗规划的核心。尽管多模态视觉-语言模型通过文本描述提升了语义理解,但在标注稀疏、设备导致图像退化的实际临床环境中,其鲁棒性仍未充分探索。我们提出BiCLIP(双向一致语言-图像处理),通过双向多模态融合机制,使视觉特征可迭代优化文本表示,实现更优语义对齐;同时引入增强一致性目标,稳定中间表示对扰动输入的响应。在QaTa-COV19与MosMedData+基准上的评估表明,BiCLIP持续优于当前主流的纯图像与多模态基线。尤其在仅使用1%标注数据训练时仍表现优异,并显著抵抗运动模糊与低剂量CT噪声等临床伪影。

原文摘要 · Abstract (English)

Medical image segmentation is a cornerstone of computer-assisted diagnosis and treatment planning. While recent multimodal vision-language models have shown promise in enhancing semantic understanding through textual descriptions, their resilience in "in-the-wild" clinical settings-characterized by scarce annotations and hardware-induced image degradations-remains under-explored. We introduce BiCLIP (Bidirectional and Consistent Language-Image Processing), a framework engineered to bolster robustness in medical segmentation. BiCLIP features a bidirectional multimodal fusion mechanism that enables visual features to iteratively refine textual representations, ensuring superior semantic alignment. To further stabilize learning, we implement an augmentation consistency objective that regularizes intermediate representations against perturbed input views. Evaluation on the QaTa-COV19 and MosMedData+ benchmarks demonstrates that BiCLIP consistently surpasses state-of-the-art image-only and multimodal baselines. Notably, BiCLIP maintains high performance when trained on as little as 1% of labeled data and exhibits significant resistance to clinical artifacts, including motion blur and low-dose CT noise.

医学图像多模态鲁棒分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。