自动标注数据集可显著提升甲状腺结节分类模型性能。
Effectiveness of Automatically Curated Dataset in Thyroid Nodules Classification Algorithms Using Deep Learning
- 用自动标注数据训练模型,替代人工标注。
- 自动数据集使模型AUC达0.694,优于人工标注的0.643。
- 全量自动数据比精选高精度子集更有效,适合临床应用。
甲状腺结节癌症诊断常依赖超声图像。已有研究显示,深度学习算法在分类良恶性结节方面可达到放射科医生水平。然而,由于数据标注耗时,深度学习模型训练数据常受限。此前研究提出一种自动标注甲状腺结节数据的方法,其生成数据的产率可达63%,准确率为83%。但该数据对深度学习模型的实际价值尚不明确。本研究通过实验验证:将深度学习模型分别在人工标注和自动标注数据集上训练,并使用准确率更高的子集探索最优使用方式。结果显示,人工标注数据训练的模型AUC为0.643(95% CI: 0.62, 0.66),显著低于自动标注数据训练的0.694(95% CI: 0.67, 0.73,P < .001)。而高精度子集训练的模型AUC为0.689(95% CI: 0.66, 0.72,P > .43),与全量自动数据无显著差异。结论表明,使用自动标注数据能大幅提升模型性能,建议使用全部数据而非仅高精度子集。
原文摘要 · Abstract (English)
The diagnosis of thyroid nodule cancers commonly utilizes ultrasound images. Several studies showed that deep learning algorithms designed to classify benign and malignant thyroid nodules could match radiologists' performance. However, data availability for training deep learning models is often limited due to the significant effort required to curate such datasets. The previous study proposed a method to curate thyroid nodule datasets automatically. It was tested to have a 63% yield rate and 83% accuracy. However, the usefulness of the generated data for training deep learning models remains unknown. In this study, we conducted experiments to determine whether using a automatically-curated dataset improves deep learning algorithms' performance. We trained deep learning models on the manually annotated and automatically-curated datasets. We also trained with a smaller subset of the automatically-curated dataset that has higher accuracy to explore the optimum usage of such dataset. As a result, the deep learning model trained on the manually selected dataset has an AUC of 0.643 (95% confidence interval [CI]: 0.62, 0.66). It is significantly lower than the AUC of the 6automatically-curated dataset trained deep learning model, 0.694 (95% confidence interval [CI]: 0.67, 0.73, P < .001). The AUC of the accurate subset trained deep learning model is 0.689 (95% confidence interval [CI]: 0.66, 0.72, P > .43), which is insignificantly worse than the AUC of the full automatically-curated dataset. In conclusion, we showed that using a automatically-curated dataset can substantially increase the performance of deep learning algorithms, and it is suggested to use all the data rather than only using the accurate subset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。