arXiv:2510.01547cs.CVcs.LG2025-10被引 1

用小数据训练的贝叶斯模型提升口腔癌诊断可靠性

Robust Classification of Oral Cancer with Limited Training Data

  • 结合CNN与贝叶斯推理,通过不确定性量化增强模型可信度
  • 小数据下对真实照片仍达88%准确率,优于传统CNN的72.94%
  • 适合医疗资源匮乏地区,可帮助基层医生做早期筛查

口腔癌是全球高发癌症之一,尤其在医疗资源不足地区死亡率较高。早期诊断对降低死亡率至关重要,但受限于健康项目少、基础设施薄弱和医护人力短缺。传统深度学习模型依赖点估计,易产生过度自信,且需大量数据以避免过拟合,这在小样本场景下难以实现。为此,本文提出一种混合模型,将卷积神经网络(CNN)与贝叶斯深度学习结合,利用变分推断进行不确定性量化,以提升可靠性。模型基于智能手机拍摄的彩色图像训练,并在三个不同测试集上评估。在分布相似的数据集上,该方法达到94%准确率,与传统CNN相当;而在真实世界复杂多变的照片数据上,其准确率达88%,显著优于传统CNN的72.94%。置信度分析显示,正确分类样本不确定性低(高置信),错误分类样本不确定性高(低置信)。结果表明,贝叶斯推断在数据稀缺环境下能有效提升模型可靠性与泛化能力,助力早期口腔癌诊断。

原文摘要 · Abstract (English)

Oral cancer ranks among the most prevalent cancers globally, with a particularly high mortality rate in regions lacking adequate healthcare access. Early diagnosis is crucial for reducing mortality; however, challenges persist due to limited oral health programs, inadequate infrastructure, and a shortage of healthcare practitioners. Conventional deep learning models, while promising, often rely on point estimates, leading to overconfidence and reduced reliability. Critically, these models require large datasets to mitigate overfitting and ensure generalizability, an unrealistic demand in settings with limited training data. To address these issues, we propose a hybrid model that combines a convolutional neural network (CNN) with Bayesian deep learning for oral cancer classification using small training sets. This approach employs variational inference to enhance reliability through uncertainty quantification. The model was trained on photographic color images captured by smartphones and evaluated on three distinct test datasets. The proposed method achieved 94% accuracy on a test dataset with a distribution similar to that of the training data, comparable to traditional CNN performance. Notably, for real-world photographic image data, despite limitations and variations differing from the training dataset, the proposed model demonstrated superior generalizability, achieving 88% accuracy on diverse datasets compared to 72.94% for traditional CNNs, even with a smaller dataset. Confidence analysis revealed that the model exhibits low uncertainty (high confidence) for correctly classified samples and high uncertainty (low confidence) for misclassified samples. These results underscore the effectiveness of Bayesian inference in data-scarce environments in enhancing early oral cancer diagnosis by improving model reliability and generalizability.

口腔癌小样本贝叶斯医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。