用高效模型与增强策略提升罕见乳腺癌亚型分类准确率
Breast Tumor Classification Using EfficientNet Deep Learning Model
- 采用EfficientNet结合数据增强和代价敏感学习缓解数据不平衡
- 多分类准确率从91.27%提升至95.04%,良性病例召回率达0.95
- 特别改善黏液性及乳头状癌等少见亚型的识别精度,适合临床辅助诊断
在组织病理图像上精确分类乳腺癌有助于显著提升肿瘤诊疗效果。医学图像数据集固有的类别不平衡问题导致某些肿瘤亚型出现频率极低,造成模型偏向多数类而忽略关键少数类。本文采用先进的EfficientNet卷积神经网络,在保持高精度与计算效率平衡的同时,引入密集数据增强与代价敏感学习策略,有效提升罕见肿瘤类型的表征能力。通过迁移学习,将二分类预训练权重用于多分类任务,增强了对BreakHis数据集复杂模式的捕捉能力。结果表明,二分类召回率从0.92提升至0.95,准确率由97.35%增至98.23%;多分类任务准确率从常规增强下的91.27%提高至密集增强下的94.54%,使用迁移学习后达95.04%。混淆矩阵分析证实,该方法显著提升了黏液性癌、乳头状癌等少数类别的精确率,同时保持了这些关键亚型的高召回率。
原文摘要 · Abstract (English)
Precise breast cancer classification on histopathological images has the potential to greatly improve the diagnosis and patient outcome in oncology. The data imbalance problem largely stems from the inherent imbalance within medical image datasets, where certain tumor subtypes may appear much less frequently. This constitutes a considerable limitation in biased model predictions that can overlook critical but rare classes. In this work, we adopted EfficientNet, a state-of-the-art convolutional neural network (CNN) model that balances high accuracy with computational cost efficiency. To address data imbalance, we introduce an intensive data augmentation pipeline and cost-sensitive learning, improving representation and ensuring that the model does not overly favor majority classes. This approach provides the ability to learn effectively from rare tumor types, improving its robustness. Additionally, we fine-tuned the model using transfer learning, where weights in the beginning trained on a binary classification task were adopted to multi-class classification, improving the capability to detect complex patterns within the BreakHis dataset. Our results underscore significant improvements in the binary classification performance, achieving an exceptional recall increase for benign cases from 0.92 to 0.95, alongside an accuracy enhancement from 97.35 % to 98.23%. Our approach improved the performance of multi-class tasks from 91.27% with regular augmentation to 94.54% with intensive augmentation, reaching 95.04% with transfer learning. This framework demonstrated substantial gains in precision in the minority classes, such as Mucinous carcinoma and Papillary carcinoma, while maintaining high recall consistently across these critical subtypes, as further confirmed by confusion matrix analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。