用自监督学习自动识别X光片解剖部位,准确率超96%。
Self-Supervised Radiograph Anatomical Region Classification -- How Clean Is Your Real-World Data?
- 用SimCLR、BYOL等自监督方法分类14类骨骼部位。
- 仅用1%标签数据就达92.2%准确率,适合低资源场景。
- 可发现35%标注错误,提升真实数据质量。
现代医学影像深度学习流程依赖于准确的解剖区域标签。确定解剖区域有助于选择下游模型并构建高质量数据集。然而外部数据常缺标签或含录入错误。我们针对48,434张骨骼X光片的内部数据集,验证了SimCLR、BYOL等自监督方法及监督对比学习的有效性,单模型线性评估准确率达96.6%,集成方法达97.7%。仅需训练集1%的标注样本即可实现92.2%准确率,适用于低标签场景。专家对最佳单模型测试集错误进行复核,发现35%标签错误和11%域外图像。修正后,单模型与集成模型理论准确率分别提升至98.0%和98.8%。
原文摘要 · Abstract (English)
Modern deep learning-based clinical imaging workflows rely on accurate labels of the examined anatomical region. Knowing the anatomical region is required to select applicable downstream models and to effectively generate cohorts of high quality data for future medical and machine learning research efforts. However, this information may not be available in externally sourced data or generally contain data entry errors. To address this problem, we show the effectiveness of self-supervised methods such as SimCLR and BYOL as well as supervised contrastive deep learning methods in assigning one of 14 anatomical region classes in our in-house dataset of 48,434 skeletal radiographs. We achieve a strong linear evaluation accuracy of 96.6% with a single model and 97.7% using an ensemble approach. Furthermore, only a few labeled instances (1% of the training set) suffice to achieve an accuracy of 92.2%, enabling usage in low-label and thus low-resource scenarios. Our model can be used to correct data entry mistakes: a follow-up analysis of the test set errors of our best-performing single model by an expert radiologist identified 35% incorrect labels and 11% out-of-domain images. When accounted for, the radiograph anatomical region labelling performance increased -- without and with an ensemble, respectively -- to a theoretical accuracy of 98.0% and 98.8%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。