用少量乳腺MRI数据即可训练出顶尖水平的癌症分类AI模型
Exploring Patient Data Requirements in Training Effective AI Models for MRI-based Breast Cancer Classification
- 基于大模型微调,大幅降低医疗数据需求
- 超50例患者后增大数据量对性能提升有限
- 简单集成策略能进一步提效,适合资源有限医院
过去十年,众多初创公司推出基于AI的临床决策支持系统。然而,医疗决策的敏感性引发对依赖外部软件的担忧,包括影像模态差异、设备异构、法律风险及对抗攻击等问题。得益于开源机器学习研究,基础模型已公开可用,使医疗机构可自主训练AI模型以规避上述风险。在此背景下,关键问题浮现:医疗机构需多少数据才能训练出有效模型?本研究聚焦乳腺癌检测这一高发疾病(约8位女性中1位患病),通过大规模实验分析不同训练集规模的影响。结果表明,只要利用基础模型,无需积累十年的MRI数据,即可训练出媲美前沿水平的模型。此外,当患者数超过50时,继续增加数据量对性能影响微乎其微;而简单的模型集成可进一步提升效果,且不增加复杂度。
原文摘要 · Abstract (English)
The past decade has witnessed a substantial increase in the number of startups and companies offering AI-based solutions for clinical decision support in medical institutions. However, the critical nature of medical decision-making raises several concerns about relying on external software. Key issues include potential variations in image modalities and the medical devices used to obtain these images, potential legal issues, and adversarial attacks. Fortunately, the open-source nature of machine learning research has made foundation models publicly available and straightforward to use for medical applications. This accessibility allows medical institutions to train their own AI-based models, thereby mitigating the aforementioned concerns. Given this context, an important question arises: how much data do medical institutions need to train effective AI models? In this study, we explore this question in relation to breast cancer detection, a particularly contested area due to the prevalence of this disease, which affects approximately 1 in every 8 women. Through large-scale experiments on various patient sizes in the training set, we show that medical institutions do not need a decade's worth of MRI images to train an AI model that performs competitively with the state-of-the-art, provided the model leverages foundation models. Furthermore, we observe that for patient counts greater than 50, the number of patients in the training set has a negligible impact on the performance of models and that simple ensembles further improve the results without additional complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。