用生成模型大量合成真实肺部肿瘤图像,解决医疗数据少难题
FreeTumor: Large-Scale Generative Tumor Synthesis in Computed Tomography Images for Improving Tumor Recognition
- 结合少量标注数据与海量无标注数据训练生成模型
- 合成超16万张带肿瘤的CT图像,数据量扩增40倍以上
- 医生辨识真假肿瘤准确率仅60.8%,适合临床数据增强
肿瘤是全球主要死因之一,每年约有1000万人死于相关疾病。人工智能驱动的肿瘤识别为更精准筛查和诊断带来可能,但受限于标注数据稀缺,需放射科医生投入大量精力标注。为此,我们提出FreeTumor——一种创新的生成式AI框架,用于大规模合成计算机断层扫描(CT)图像中的肿瘤,缓解数据短缺问题。FreeTumor有效利用有限标注数据与大规模无标注数据进行训练,通过释放海量数据潜力,实现大量逼真肿瘤的合成,以扩充训练集。我们从33个来源整合了161,310份公开可得的CT体数据,其中仅2.3%包含标注肿瘤。为验证合成肿瘤的真实性,我们邀请13位持证放射科医生参与视觉图灵测试,结果表明其辨别合成与真实肿瘤的敏感度仅为51.1%,准确率仅60.8%。得益于高质量合成,FreeTumor使识别训练数据规模扩大超过40倍,在性能上显著优于当前主流合成方法及基础模型。这些成果预示其在临床应用中的广阔前景,有望提升肿瘤治疗效果并改善患者生存率。
原文摘要 · Abstract (English)
Tumor is a leading cause of death worldwide, with an estimated 10 million deaths attributed to tumor-related diseases every year. AI-driven tumor recognition unlocks new possibilities for more precise and intelligent tumor screening and diagnosis. However, the progress is heavily hampered by the scarcity of annotated datasets, which demands extensive annotation efforts by radiologists. To tackle this challenge, we introduce FreeTumor, an innovative Generative AI (GAI) framework to enable large-scale tumor synthesis for mitigating data scarcity. Specifically, FreeTumor effectively leverages a combination of limited labeled data and large-scale unlabeled data for tumor synthesis training. Unleashing the power of large-scale data, FreeTumor is capable of synthesizing a large number of realistic tumors on images for augmenting training datasets. To this end, we create the largest training dataset for tumor synthesis and recognition by curating 161,310 publicly available Computed Tomography (CT) volumes from 33 sources, with only 2.3% containing annotated tumors. To validate the fidelity of synthetic tumors, we engaged 13 board-certified radiologists in a Visual Turing Test to discern between synthetic and real tumors. Rigorous clinician evaluation validates the high quality of our synthetic tumors, as they achieved only 51.1% sensitivity and 60.8% accuracy in distinguishing our synthetic tumors from real ones. Through high-quality tumor synthesis, FreeTumor scales up the recognition training datasets by over 40 times, showcasing a notable superiority over state-of-the-art AI methods including various synthesis methods and foundation models. These findings indicate promising prospects of FreeTumor in clinical applications, potentially advancing tumor treatments and improving the survival rates of patients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。