arXiv:2508.12108cs.CV2025-08被引 1

用少量3DCT与报告对训练医学影像视觉语言模型,提升下游任务表现。

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

  • 融合单模态自监督学习与分层对比学习,增强多模态表征能力。
  • 仅用38,875组数据实现分割、问答等任务的先进性能。
  • 适合医疗影像研究者,尤其关注小样本医学多模态建模的人群。

视觉语言模型(VLM)在医学领域受到越来越多关注,尤其在通用领域成功案例如CLIP之后。然而,与二维图像和文本的简单配对不同,获取三维医学影像(如CT扫描)与对应放射科报告的大规模成对数据仍具挑战性且耗时。为应对这一难题,我们提出一种新型视觉语言预训练框架VELVET-Med,专为有限的体积数据(如3D CT和相应放射科报告)设计。本方法不依赖大规模数据收集,而是聚焦于有效预训练目标与模型架构的设计。主要贡献包括:1)将单模态自监督学习引入视觉语言预训练框架,该方向在现有文献中常被忽视;2)提出新型语言编码器TriBERT,用于学习多层次文本语义;3)设计分层对比学习机制,捕捉多层级视觉-语言对应关系。仅使用38,875组扫描-报告配对,我们的方法旨在揭示体积医学图像与临床叙述中蕴含的丰富空间与语义关系,从而增强编码器的泛化能力。所获得的编码器展现出强迁移性,在多种下游任务中达到领先水平,包括3D分割、跨模态检索、视觉问答和报告生成。

原文摘要 · Abstract (English)

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating large-scale paired data in the medical field for volumetric modalities such as CT scans remains a challenging and time-intensive process. This difficulty often limits the performance on downstream tasks. To address these challenges, we propose a novel vision-language pre-training (VLP) framework, termed as \textbf{VELVET-Med}, specifically designed for limited volumetric data such as 3D CT and associated radiology reports. Instead of relying on large-scale data collection, our method focuses on the development of effective pre-training objectives and model architectures. The key contributions are: 1) We incorporate uni-modal self-supervised learning into VLP framework, which are often underexplored in the existing literature. 2) We propose a novel language encoder, termed as \textbf{TriBERT}, for learning multi-level textual semantics. 3) We devise the hierarchical contrastive learning to capture multi-level vision-language correspondence. Using only 38,875 scan-report pairs, our approach seeks to uncover rich spatial and semantic relationships embedded in volumetric medical images and corresponding clinical narratives, thereby enhancing the generalization ability of the learned encoders. The resulting encoders exhibit strong transferability, achieving state-of-the-art performance across a wide range of downstream tasks, including 3D segmentation, cross-modal retrieval, visual question answering, and report generation.

医学影像多模态小样本学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。