arXiv:2508.06453cs.CVcs.AI2025-08被引 1

将报告文本融入视觉模型,提升CT病灶分割精度

Text Embedded Swin-UMamba for DeepLesion Segmentation

  • 用文本嵌入增强Swin-UMamba架构,融合影像与报告信息
  • 测试集上达82.64的Dice分数,95%置信区间下显著优于基线
  • 适合医学影像分析、多模态模型研究者参考

CT图像中病灶的分割可实现慢性疾病(如淋巴瘤)的自动评估。将大语言模型(LLM)引入病灶分割流程,有望结合影像特征与放射科报告中的病灶描述。本研究探索在Swin-UMamba架构中整合文本信息用于病灶分割的可行性。使用公开的ULS23 DeepLesion数据集及报告中的简短病灶描述。在测试集上,本方法获得82.64的高Dice分数和6.34像素的低豪斯多夫距离。所提出的Text-Swin-U/Mamba模型显著优于先前方法:较基于LLM的LanGuideMedSeg模型提升37.79%(p < 0.001),分别超过纯图像的XLSTM-UNet和nnUNet模型2.58%和1.01%。代码与数据集见https://github.com/ruida/LLM-Swin-UMamba。

原文摘要 · Abstract (English)

Segmentation of lesions on CT enables automatic measurement for clinical assessment of chronic diseases (e.g., lymphoma). Integrating large language models (LLMs) into the lesion segmentation workflow has the potential to combine imaging features with descriptions of lesion characteristics from the radiology reports. In this study, we investigate the feasibility of integrating text into the Swin-UMamba architecture for the task of lesion segmentation. The publicly available ULS23 DeepLesion dataset was used along with short-form descriptions of the findings from the reports. On the test dataset, our method achieved a high Dice score of 82.64, and a low Hausdorff distance of 6.34 pixels was obtained for lesion segmentation. The proposed Text-Swin-U/Mamba model outperformed prior approaches: 37.79% improvement over the LLM-driven LanGuideMedSeg model (p < 0.001), and surpassed the purely image-based XLSTM-UNet and nnUNet models by 2.58% and 1.01%, respectively. The dataset and code can be accessed at https://github.com/ruida/LLM-Swin-UMamba

病灶分割多模态医学影像文本嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。