arXiv:2412.13558eess.IVcs.CL2024-12被引 22

模仿放射科医生读片流程,提升3D医学影像理解能力

Read Like a Radiologist: Efficient Vision-Language Model for 3D Medical Imaging Interpretation

  • 用2D Transformer逐切片分析,再整合跨切片信息
  • 在胸部CT和直肠MRI上生成更连贯、临床相关的报告
  • 支持任意切片长度与多平面数据,通用性强

近期医学视觉语言模型(VLM)在二维医学图像解读中表现良好,但扩展到三维医学影像仍面临计算复杂性和数据稀缺的挑战。现有少数针对3D医学影像的VLM仅能将三维图像表示为子体积特征集合,导致沿z轴特征高度相关,忽略切片特有的临床细节,尤其在相邻切片冗余度低的3D影像中更为明显。为此,我们提出MS-VLM,模拟放射科医生的读片流程:逐切片分析并综合多切片信息。该模型利用自监督2D Transformer编码器,从一系列切片特征求得具有切片间依赖关系的三维表示。不受子体积分块限制,MS-VLM可处理任意切片长度的三维影像,并兼容不同平面与阶段获取的多模态数据。在公开的胸部CT数据集CT-RATE和内部直肠MRI数据集上评估,其在放射科报告生成任务中均优于现有方法,生成报告更具连贯性与临床相关性。结果表明,MS-VLM有望推动三维医学影像理解发展,增强医学VLM的鲁棒性。

原文摘要 · Abstract (English)

Recent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent VLMs specified for 3D medical imaging have emerged, all are limited to learning volumetric representation of a 3D medical image as a set of sub-volumetric features. Such process introduces overly correlated representations along the z-axis that neglect slice-specific clinical details, particularly for 3D medical images where adjacent slices have low redundancy. To address this limitation, we introduce MS-VLM that mimic radiologists' workflow in 3D medical image interpretation. Specifically, radiologists analyze 3D medical images by examining individual slices sequentially and synthesizing information across slices and views. Likewise, MS-VLM leverages self-supervised 2D transformer encoders to learn a volumetric representation that capture inter-slice dependencies from a sequence of slice-specific features. Unbound by sub-volumetric patchification, MS-VLM is capable of obtaining useful volumetric representations from 3D medical images with any slice length and from multiple images acquired from different planes and phases. We evaluate MS-VLM on publicly available chest CT dataset CT-RATE and in-house rectal MRI dataset. In both scenarios, MS-VLM surpasses existing methods in radiology report generation, producing more coherent and clinically relevant reports. These findings highlight the potential of MS-VLM to advance 3D medical image interpretation and improve the robustness of medical VLMs.

3D医学影像视觉语言模型放射科读片自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。