用AI和众包构建高质量医学图像分割训练集
Coupling AI and Citizen Science in Creation of Enhanced Training Dataset for Medical Image Segmentation
- 结合众包标注与MedSAM AI,高效生成专家级标注
- 引入pix2pixGAN生成真实形态的合成图像,数据量提升3倍
- 适合数据少时提升模型性能,尤其适合医疗AI研究者
近年来,医学影像与人工智能的进步显著提升了诊断能力,但深度学习模型的有效开发仍受限于高质量标注数据集的缺乏。传统由医疗专家进行的手动标注耗时且资源消耗大,制约了数据集的可扩展性。本文提出一种鲁棒且通用的框架,融合AI与众包,提升多模态医学图像数据集的质量与数量。通过一个用户友好的在线平台,使多样化的众包标注者能高效标注医学图像。结合MedSAM分割AI模型,利用算法融合众包标注结果,在加速标注的同时保持专家级质量。此外,采用pix2pixGAN生成对抗网络,生成具有真实解剖形态特征的合成图像以扩充训练数据。上述方法整合为统一框架,形成增强型数据集,可作为通用预处理流程,显著提升任意医学深度学习分割模型的训练效果。实验表明,该框架在训练数据有限时能显著提升模型性能。
原文摘要 · Abstract (English)
Recent advancements in medical imaging and artificial intelligence (AI) have greatly enhanced diagnostic capabilities, but the development of effective deep learning (DL) models is still constrained by the lack of high-quality annotated datasets. The traditional manual annotation process by medical experts is time- and resource-intensive, limiting the scalability of these datasets. In this work, we introduce a robust and versatile framework that combines AI and crowdsourcing to improve both the quality and quantity of medical image datasets across different modalities. Our approach utilises a user-friendly online platform that enables a diverse group of crowd annotators to label medical images efficiently. By integrating the MedSAM segmentation AI with this platform, we accelerate the annotation process while maintaining expert-level quality through an algorithm that merges crowd-labelled images. Additionally, we employ pix2pixGAN, a generative AI model, to expand the training dataset with synthetic images that capture realistic morphological features. These methods are combined into a cohesive framework designed to produce an enhanced dataset, which can serve as a universal pre-processing pipeline to boost the training of any medical deep learning segmentation model. Our results demonstrate that this framework significantly improves model performance, especially when training data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。