对比视觉与统计评估方法,提出深度学习图像分析的可靠验证方案。
Performance evaluation of deep learning models for image analysis: considerations for visual control and statistical metrics
- 结合视觉检查与统计指标评估模型性能
- 强调高质量标注数据与独立测试集的重要性
- 适合病理学、药物安全评估等需高可靠性场景
基于深度学习的自动化图像分析(DL-AIA)在特征量化任务中已超越训练过的病理科医生。其应用正从原理验证扩展至临床诊断病理、毒理病理安全性评估及重复性研究。为确保应用的安全可靠,必须进行全面客观的泛化性能评估(即算法对目标模式的准确预测能力),并可能评估模型鲁棒性(即在不同来源图像上保持预测精度的能力)。本文回顾兽医病理学文献中的性能评估实践,识别出两种方法:1)仅依赖视觉性能控制(即人工审阅算法预测结果)并辅以次级性能指标验证;2)结合统计性能控制,需在模型训练前构建数据集并分离保留测试集。本文比较了两类方法的优劣,并讨论了严格统计评估的关键考量,包括指标选择、测试集图像构成、真实标签质量、重抽样方法(如自助法)、多模型统计比较及模型稳定性评估。结论认为,视觉与统计评估互补,二者结合能最深入揭示模型性能与误差来源。
原文摘要 · Abstract (English)
Deep learning-based automated image analysis (DL-AIA) has been shown to outperform trained pathologists in tasks related to feature quantification. Related to these capacities the use of DL-AIA tools is currently extending from proof-of-principle studies to routine applications such as patient samples (diagnostic pathology), regulatory safety assessment (toxicologic pathology), and recurrent research tasks. To ensure that DL-AIA applications are safe and reliable, it is critical to conduct a thorough and objective generalization performance assessment (i.e., the ability of the algorithm to accurately predict patterns of interest) and possibly evaluate model robustness (i.e., the algorithm's capacity to maintain predictive accuracy on images from different sources). In this article, we review the practices for performance assessment in veterinary pathology publications by which two approaches were identified: 1) Exclusive visual performance control (i.e. eyeballing of algorithmic predictions) plus validation of the models application utilizing secondary performance indices, and 2) Statistical performance control (alongside the other methods), which requires a dataset creation and separation of an hold-out test set prior to model training. This article compares the strengths and weaknesses of statistical and visual performance control methods. Furthermore, we discuss relevant considerations for rigorous statistical performance evaluation including metric selection, test dataset image composition, ground truth label quality, resampling methods such as bootstrapping, statistical comparison of multiple models, and evaluation of model stability. It is our conclusion that visual and statistical evaluation have complementary strength and a combination of both provides the greatest insight into the DL model's performance and sources of error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。