MedVista3D用多尺度视觉语言模型提升3D CT诊断准确性,减少漏诊和报告不一致。
MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting
- 融合局部与全局信息,实现3D CT中病灶精确定位与整体语义理解。
- 在零样本疾病分类、报告检索等任务上达当前最佳,跨任务迁移能力强。
- 适合医学影像分析、AI辅助诊断系统研发者使用,尤其关注临床实用性。
放射科诊断错误——如漏诊、注意力盲区和沟通失误——在临床实践中仍普遍存在。这些问题常源于局部异常的遗漏、全局上下文不足以及报告语言的不一致性。这些挑战在3D影像中尤为突出,医生需审查每例扫描数百张切片。解决此问题需具备精确的局部检测、全局体积推理及语义一致的自然语言报告能力。然而,现有3D视觉-语言模型难以同时满足三方面需求,缺乏空间推理所需的局部-全局理解,且对未校准报告中的变异性和噪声处理能力弱。我们提出MedVista3D,一种用于3D CT分析的多尺度语义增强型视觉-语言预训练框架。为实现疾病检测与整体解读的联合优化,该模型在全体积上下文中进行局部与全局图像-文本对齐,以学习细粒度表征。为应对报告差异性,引入语言模型重写与放射科语义匹配库,实现语义感知对齐。MedVista3D在零样本疾病分类、报告检索和医学视觉问答任务上达到最新水平,并在器官分割与预后预测任务中表现出良好迁移能力。代码与数据集将公开。
原文摘要 · Abstract (English)
Radiologic diagnostic errors-under-reading errors, inattentional blindness, and communication failures-remain prevalent in clinical practice. These issues often stem from missed localized abnormalities, limited global context, and variability in report language. These challenges are amplified in 3D imaging, where clinicians must examine hundreds of slices per scan. Addressing them requires systems with precise localized detection, global volume-level reasoning, and semantically consistent natural language reporting. However, existing 3D vision-language models are unable to meet all three needs jointly, lacking local-global understanding for spatial reasoning and struggling with the variability and noise of uncurated radiology reports. We present MedVista3D, a multi-scale semantic-enriched vision-language pretraining framework for 3D CT analysis. To enable joint disease detection and holistic interpretation, MedVista3D performs local and global image-text alignment for fine-grained representation learning within full-volume context. To address report variability, we apply language model rewrites and introduce a Radiology Semantic Matching Bank for semantics-aware alignment. MedVista3D achieves state-of-the-art performance on zero-shot disease classification, report retrieval, and medical visual question answering, while transferring well to organ segmentation and prognosis prediction. Code and datasets will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。