通过2000+训练组合对比主流CNN模型,提升乳腺癌病理图像分类准确率。
Comparative Analysis and Ensemble Enhancement of Leading CNN Architectures for Breast Cancer Classification
- 系统比较多种CNN模型,优化数据增强与全连接层设计。
- 在BreakHis数据集上达到99.75%准确率,Bach数据集达95.18%。
- 提出三模型集成架构,适合医学图像分类研究者参考。
本研究提出一种新颖且精确的乳腺癌分类方法,基于组织病理图像进行系统性分析。通过在不同图像数据集上对比主流卷积神经网络(CNN)模型,确定其最优超参数,并按分类效能进行排名。为提升各模型性能,探索了数据增强、替代全连接层、训练超参数设置及微调权重与使用预训练权重的优劣。方法包含多个原创设计,如序列化生成数据集以确保训练条件一致,显著缩短训练时间;结合自动化结果整理,实现了超过2000种训练组合的探索,这是迄今前所未有的全面比较。研究明确了实现高精度独立CNN模型所需的配置,并据此提出集成架构:将三个高性能独立模型与多样化分类器堆叠,进一步提升分类准确率。该方法在BreakHis x40和x200数据集上分别达到99.75%准确率,在Bach数据集上达95.18%。在Bach Online盲测挑战中取得89%准确率。尽管聚焦乳腺癌病理图像,但该方法同样适用于其他医学图像数据集。
原文摘要 · Abstract (English)
This study introduces a novel and accurate approach to breast cancer classification using histopathology images. It systematically compares leading Convolutional Neural Network (CNN) models across varying image datasets, identifies their optimal hyperparameters, and ranks them based on classification efficacy. To maximize classification accuracy for each model we explore, the effects of data augmentation, alternative fully-connected layers, model training hyperparameter settings, and, the advantages of retraining models versus using pre-trained weights. Our methodology includes several original concepts, including serializing generated datasets to ensure consistent data conditions across training runs and significantly reducing training duration. Combined with automated curation of results, this enabled the exploration of over 2,000 training permutations -- such a comprehensive comparison is as yet unprecedented. Our findings establish the settings required to achieve exceptional classification accuracy for standalone CNN models and rank them by model efficacy. Based on these results, we propose ensemble architectures that stack three high-performing standalone CNN models together with diverse classifiers, resulting in improved classification accuracy. The ability to systematically run so many model permutations to get the best outcomes gives rise to very high quality results, including 99.75% for BreakHis x40 and BreakHis x200 and 95.18% for the Bach datasets when split into train, validation and test datasets. The Bach Online blind challenge, yielded 89% using this approach. Whilst this study is based on breast cancer histopathology image datasets, the methodology is equally applicable to other medical image datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。