arXiv:2508.13796cs.CVcs.AI2025-08ICCV被引 2

用报告和图像联合训练,让癌症分割模型既准又说得清。

A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports

  • 双分支视觉编码器融合ViT与Swin,结合临床报告做跨模态注意力。
  • 在乳腺超声数据集上达99%骰子系数,比U-Net等模型更优。
  • 自动生成分割图、置信度图和诊断理由,适合临床可信辅助。

我们提出Med-CTX,一种全Transformer架构的多模态可解释性乳腺癌超声分割框架。通过融合临床放射科报告提升性能与可解释性。该方法采用双分支视觉编码器(结合ViT与Swin Transformer)及不确定性感知融合机制,利用BioClinicalBERT对具有BI-RADS语义的临床文本编码,并通过跨模态注意力与视觉特征融合,生成符合临床逻辑的模型解释。该框架同时输出分割掩码、不确定性图和诊断推理,增强计算机辅助诊断的可信度与透明性。在BUS-BRA数据集上,Med-CTX取得99%的骰子系数和95%的交并比,优于现有基线模型(U-Net、ViT、Swin)。消融实验表明,临床文本缺失会导致骰子系数下降5.4%、CIDEr下降31%。模型实现良好多模态对齐(CLIP得分85%)与更优的置信度校准(ECE: 3.2%),为可信多模态医疗架构树立新标准。

原文摘要 · Abstract (English)

We introduce Med-CTX, a fully transformer based multimodal framework for explainable breast cancer ultrasound segmentation. We integrate clinical radiology reports to boost both performance and interpretability. Med-CTX achieves exact lesion delineation by using a dual-branch visual encoder that combines ViT and Swin transformers, as well as uncertainty aware fusion. Clinical language structured with BI-RADS semantics is encoded by BioClinicalBERT and combined with visual features utilising cross-modal attention, allowing the model to provide clinically grounded, model generated explanations. Our methodology generates segmentation masks, uncertainty maps, and diagnostic rationales all at once, increasing confidence and transparency in computer assisted diagnosis. On the BUS-BRA dataset, Med-CTX achieves a Dice score of 99% and an IoU of 95%, beating existing baselines U-Net, ViT, and Swin. Clinical text plays a key role in segmentation accuracy and explanation quality, as evidenced by ablation studies that show a -5.4% decline in Dice score and -31% in CIDEr. Med-CTX achieves good multimodal alignment (CLIP score: 85%) and increased confi dence calibration (ECE: 3.2%), setting a new bar for trustworthy, multimodal medical architecture.

医学图像可解释性多模态分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。