提升眼底OCT分类可信度,让模型判断更安全可靠。
Calibrated Hybrid CNN-Transformer for Retinal OCT Classification

- 混合CNN-Transformer架构+梯度提升分类器,兼顾精度与可解释性。
- 在84,495张扫描图像上达到95.4%准确率,校准误差降低12倍。
- 首次同时验证三重临床安全机制,适合医疗部署场景。
用于视网膜光学相干断层扫描(OCT)分类的深度模型虽报告高准确率,却极少评估其置信度是否可信——这在错误但自信的诊断可能延误救命治疗时至关重要。本文采用混合卷积-Transformer编码器,搭配梯度提升(XGBoost)分类头,并集成三重临床安全层:置信度校准、分布外(OOD)样本拒绝及逐预测不确定性标记。在四类OCT数据集(84,495次扫描)上,模型准确率达95.4%,校准误差降低十二倍(预期校准误差,ECE = 0.0024),表明其置信度与其真实准确性高度一致。据我们所知,这是首个联合验证三项安全机制的OCT分类器,提供公开权重并支持多种子可复现评估。
原文摘要 · Abstract (English)
Deep models for retinal optical coherence tomography (OCT) classification report high accuracy but rarely report whether their confidence can be trusted -- a gap that matters when a wrong-but-confident reading delays sight-saving treatment. We pair a hybrid convolutional-Transformer encoder with a gradient-boosting (XGBoost) classification head and a three-part clinical safety layer: confidence calibration, out-of-distribution (OOD) rejection, and per-prediction uncertainty flagging. On four-class OCT (84,495 scans) the model reaches 95.4% accuracy while cutting calibration error twelve-fold (expected calibration error, ECE = 0.0024), so the confidence it reports tracks its true accuracy. To our knowledge this is the first OCT classifier to validate all three safety mechanisms jointly, with public weights and reproducible multi-seed evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。