arXiv:2607.12075eess.IVcs.AI2026-07

基于深度集成的甲状腺结节超声分类模型,实现精准风险分层与可信赖的自动筛选。

Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation

  • 采用五成员深度集成结合校准机制,通过注意力模块提升分类性能。
  • 内部验证中准确率超98%,外部数据集下仍保持99.8%恶性肿瘤检出率。
  • 适合临床辅助决策,但需本地化校准后才能部署使用。

背景:深度学习可对超声图像中的甲状腺结节进行分类,但可靠的临床辅助决策还需具备校准概率、不确定性估计和选择性转诊能力,尤其是在数据分布漂移情况下。方法:我们构建了一个基于区域感兴趣(ROI)的五成员确定性深度集成模型,用于甲状腺结节分类与选择性图像分诊。使用TN5000数据集进行模型训练、五折交叉验证、成员向量缩放校准及折间阈值选择;以独立的外部数据集TN3K评估数据分布漂移下的表现。模型采用ConvNeXt-Tiny结构并引入挤压-激励注意力机制,利用集成均值恶性概率与互信息(MI)作为集成分歧度量,并制定三阶梯策略:无穿刺建议、细针穿刺推荐或放射科医生复核。结果:在合并的外部折叠预测中,模型达到AUC-ROC 0.9395,AP 0.9715,ECE 0.0088,Brier得分0.0813。当保留50%名义互信息时,7.2%病例被归为无穿刺建议,39.9%为穿刺推荐,52.9%进入医生复核,其中无穿刺路径阴性预测值达98.3%,恶性肿瘤捕获率99.83%。在TN3K上,AUC-ROC下降至0.7870,AP降至0.7254,ECE升至0.1899,Brier得分增至0.2281。固定使用TN5000的策略导致83.7%进入复核,1.0%进入无穿刺路径,15.3%进入穿刺推荐,无恶性病灶误入无穿刺路径,但穿刺推荐阳性预测值降至76.6%。结论:该框架在内部表现出强区分能力和良好校准性,但外部阈值迁移能力有限。选择性预测有助于识别不适合自动化分诊的图像,但在实际部署前需进行本地再校准、阈值验证及前瞻性临床评估。

原文摘要 · Abstract (English)

Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift. Methods: We developed a calibrated deterministic five-member deep ensemble for ROI-based thyroid nodule classification and selective image-based triage. TN5000 was used for model development, five-fold cross-validation, member-wise vector-scaling calibration, and fold-specific threshold selection. TN3K served as an independent external dataset-shift evaluation. The framework used ConvNeXt-Tiny with squeeze-and-excitation attention, ensemble-mean malignancy probability, and mutual information (MI) as an ensemble-disagreement score. A three-tier policy assigned images to No-FNA suggestion, FNA recommendation, or radiologist review. Results: On pooled out-of-fold TN5000 predictions, the ensemble achieved AUC-ROC 0.9395, AP 0.9715, ECE 0.0088, and Brier score 0.0813. At 50% nominal MI retention, 7.2% of cases received a No-FNA suggestion, 39.9% an FNA recommendation, and 52.9% radiologist review, with 98.3% No-FNA NPV and 99.83% malignancy capture. On TN3K, AUC-ROC decreased to 0.7870, AP to 0.7254, ECE increased to 0.1899, and Brier score to 0.2281. The frozen TN5000 policy assigned 83.7% to review, 1.0% to No-FNA, and 15.3% to FNA recommendation. No malignant image entered the No-FNA pathway, but FNA-recommendation PPV fell to 76.6%. Conclusion: The framework showed strong internal discrimination and calibration, but limited external threshold transportability. Selective prediction may help identify images unsuitable for automated triage, but local recalibration, threshold validation, and prospective clinical evaluation are required before deployment.

甲状腺结节深度集成选择性预测超声诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。