arXiv:2511.02996cs.CV2025-11被引 1

用空间语义增强医学影像的跨模态预训练,提升报告生成与诊断准确率。

SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics

  • 引入体数据空间语义和放射学知识图谱,实现结构一致的跨模态对齐
  • 在CT报告检索上比现有方法高4.3倍,异常分类提升10个百分点
  • 零样本迁移表现稳定,适合医疗影像多任务应用

视觉语言模型(VLMs)虽具强大跨模态能力,但多数研究局限于二维数据,且依赖二元监督(正负样本对),忽视了如CT扫描中连续而有序的空间依赖关系。现有方法常将体数据视为独立2D切片,损害空间连贯性并忽略丰富的临床语义。本文提出SCALE-VLP,一种软加权对比式体数据视觉语言预训练框架,融合(i)体数据空间语义以保留解剖结构,(ii)领域感知、知识注入的语义(如放射学本体)引导对齐。该方法在有限监督下生成结构一致且语义扎实的表征,展现出优异的跨任务迁移能力(检索、报告生成、分类)及跨域泛化性,无需微调即获持续提升。相比前人最优模型,SCALE-VLP在顶1报告检索上最高提升4.3倍,异常分类准确率提高10个百分点,报告生成达到ROUGE-L 0.44与BERT-F1 0.89。在外部域零样本评估中亦保持一致增益,验证其跨任务与跨域泛化能力。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated strong cross-modal capabilities, yet most work remains limited to 2D data and assumes binary supervision (i.e., positive vs. negative pairs), overlooking the continuous and structured dependencies present in volumetric data such as CT. Existing approaches often treat volumetric scans as independent 2D slices, compromising spatial coherence and underutilizing rich clinical semantics. We propose SCALE-VLP, a soft-weighted contrastive vision-language pre-training framework that integrates (i) volumetric spatial semantics to preserve anatomical structure and (ii) domain-aware, knowledge-infused semantics (e.g., radiological ontologies) to guide alignment. This yields structurally consistent and semantically grounded representations under limited supervision, demonstrating strong cross-task transferability (retrieval, report generation, and classification), and cross-domain generalizability with consistent gains without further fine-tuning. In particular, compared to the previous state of the art, SCALE-VLP achieves up to 4.3x higher top-1 CT-report retrieval, improves abnormality classification by 10 points, and reaches ROUGE-L 0.44 and BERT-F1 0.89 for report generation. Further, in zero-shot evaluation on an out-of-domain external dataset, we observe consistent gains, indicating the cross-task and cross-domain generalization ability of SCALE-VLP.

医学影像视觉语言跨模态三维建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。