用精选数据训练肺部X光大模型,省时省力还更准。
A data- and compute-efficient chest X-ray foundation model beyond aggressive scaling
- 挑出最有信息量的22.7%影像训练,避免冗余和偏差
- 仅用27.3%算力就达到甚至超过全量数据模型性能
- 适合资源有限但需高泛化能力的医疗视觉任务
医学影像大模型通常依赖海量数据预训练,但大规模数据常含大量冗余与严重类别不平衡,导致表征学习偏向常见模式;同时对数据质量差异不加区分的训练造成显著计算浪费。本文提出CheXficient,一种针对胸部X光(CXR)的高效基础模型,通过主动、有原则地筛选训练样本,仅使用1,235,004张配对CXR图像中的22.7%进行预训练,计算开销低于总预算的27.3%。在涵盖5类任务共20个基准的评估中,该模型在零样本分类、跨模态检索及疾病预测、语义分割、报告生成等下游任务上表现相当或更优。分析显示,模型系统性优先选择低频样本,提升了长尾或罕见病种的泛化能力。本工作为医疗视觉-语言模型的高效预训练与适配提供了实用路径。
原文摘要 · Abstract (English)
Foundation models for medical imaging are typically pretrained on increasingly large datasets, following a "scale-at-all-costs" paradigm. However, this strategy faces two critical challenges: large-scale medical datasets often contain substantial redundancy and severe class imbalance that bias representation learning toward over-represented patterns, and indiscriminate training regardless of heterogeneity in data quality incurs considerable computational inefficiency. Here we demonstrate that active, principled data curation during pretraining can serve as a viable, cost-effective alternative to brute-force dataset enlargement. We introduce CheXficient, a chest X-ray (CXR) foundation model that selectively prioritizes informative training samples. CheXficient is pretrained on only 22.7% of 1,235,004 paired CXR images and reports while consuming under 27.3% of the total compute budget, yet achieving comparable or superior performance to its full-data counterpart and other large-scale pretrained models. We assess CheXficient across 20 individual benchmarks spanning 5 task types, including non-adapted off-the-shelf evaluations (zero-shot findings classification and crossmodal retrieval) and adapted downstream tasks (disease prediction, semantic segmentation, and radiology report generation). Further analyses show that CheXficient systematically prioritizes under-represented training samples, improving generalizability on long-tailed or rare conditions. Overall, our work offers practical insights into the data and computation demands for efficient pretraining and downstream adaptation of medical vision-language foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。