用诊断报告生成结构化监督信号,实现轻量高效的医学影像多模态模型。
GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training
- 以诊断报告为核心构建结构化监督信号,指导模型学习
- 仅需24 GPU小时训练,零样本下CT-RATE AUC达84.8
- 适合资源受限医院本地微调,支持跨机构迁移
放射科基础模型(RFMs)沿袭了自然图像视觉-语言预训练的规模优先范式,但在3D医学影像中难以部署,因数据集较小、报告格式不一,且受隐私与算力限制需本地适配。本文提出以常规放射科报告作为可审计的诊断监督信号,驱动图像编码器、文本编码器、对齐空间及本地适配流程的学习。开发GreenRFM框架,基于四个实证原则:更精炼(More distilled)、更普遍(Ubiquitous)、语义强化(Semantic-enforcing)、任务对齐(Task-aligning),将噪声报告转化为结构化诊断信号,用于学习判别性单模态编码器和面向诊断的图文对齐空间。该模型在单张24GB GPU上仅需24 GPU小时(轻量版仅需6GB显存、4~6小时),零样本测试下达到CT-RATE AUC 84.8。在六家机构、两种模态超过20万例影像上验证,表现良好,支持向私有临床队列及骨骼肌MRI迁移。在本地机构队列中,低成本微调使宏AUC从70.5提升至82.1。对肝细胞癌微血管侵犯预测与经动脉化疗栓塞反应分析,优于现有临床评分。结果表明,以监督为中心的预训练是实现资源高效、可本地适配、以诊断为中心的医学视觉-语言表示的有效路径。
原文摘要 · Abstract (English)
Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is difficult to deploy in 3D radiology, where training corpora are smaller, reports vary across institutions, and receiving hospitals often need local adaptation under privacy and compute constraints. We ask whether routine radiology reports can instead be converted into auditable diagnostic supervision that shapes the image encoder, text encoder, aligned space, and local-adaptation procedure. We develop GreenRFM, a supervision-centric pre-training framework organized around four empirical principles: More distilled, Ubiquitous, Semantic-enforcing, and Task-aligning (MUST) supervision. These principles convert noisy reports into structured diagnostic signals and use them to learn discriminative unimodal encoders plus an aligned image--text space for diagnosis-centered multimodal use. GreenRFM requires 24 GPU-hours on a single 24GB GPU (lightweight variant: 6GB VRAM, 4~hours) and reaches a zero-shot CT-RATE AUC of 84.8. Evaluations using more than 200,000 volumes from six institutions and two modalities show transfer to private clinical cohorts and to musculoskeletal MRI. On a local institutional cohort, computationally feasible retraining raises macro-AUC from 70.5 to 82.1. The aligned space also improves hepatocellular-carcinoma microvascular-invasion prediction and trans-arterial chemoembolization response analysis over established clinical scores. These results support supervision-centric pre-training as a practical route to resource-efficient, locally adaptable, diagnosis-centered radiology vision--language representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。