arXiv:2512.00084cs.CVcs.LG2025-12中稿 · ICCV

用医学文本提升扩散模型分割效率,减少标注依赖。

A Fast and Efficient Modern BERT based Text-Conditioned Diffusion Model for Medical Image Segmentation

  • 用ModernBERT融合临床文本与图像,实现跨模态语义对齐
  • 在MIMIC-III/IV上训练,分割精度优于传统扩散模型
  • 支持快速训练,适合医疗影像分析场景

近期,去噪扩散概率模型(DPMs)在医学图像生成、去噪及下游分割任务中表现优异,但其性能受限于密集像素级标注的高成本。本文提出FastTextDiff,一种基于ModernBERT的标签高效扩散分割模型,通过整合医学文本注释增强图像语义表征。ModernBERT可处理长篇临床记录,将文本信息与图像内容紧密关联,利用MIMIC-III和MIMIC-IV数据集训练,建立视觉与文本特征间的跨模态注意力机制。该方法取代传统的Clinical BioBERT,借助FlashAttention 2与2万亿词语料库,在保持高精度的同时显著提升训练效率,验证了ModernBERT作为快速可扩展替代方案的潜力,凸显多模态技术在医学影像分析中的前景。

原文摘要 · Abstract (English)

In recent times, denoising diffusion probabilistic models (DPMs) have proven effective for medical image generation and denoising, and as representation learners for downstream segmentation. However, segmentation performance is limited by the need for dense pixel-wise labels, which are expensive, time-consuming, and require expert knowledge. We propose FastTextDiff, a label-efficient diffusion-based segmentation model that integrates medical text annotations to enhance semantic representations. Our approach uses ModernBERT, a transformer capable of processing long clinical notes, to tightly link textual annotations with semantic content in medical images. Trained on MIMIC-III and MIMIC-IV, ModernBERT encodes clinical knowledge that guides cross-modal attention between visual and textual features. This study validates ModernBERT as a fast, scalable alternative to Clinical BioBERT in diffusion-based segmentation pipelines and highlights the promise of multi-modal techniques for medical image analysis. By replacing Clinical BioBERT with ModernBERT, FastTextDiff benefits from FlashAttention 2, an alternating attention mechanism, and a 2-trillion-token corpus, improving both segmentation accuracy and training efficiency over traditional diffusion-based models.

医学图像扩散模型多模态文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。