arXiv:2601.00990eess.IVcs.CV2026-01综述

为胎儿超声图像分类设计可校准不确定性的解释性AI,提升临床可信度。

Uncertainty-Calibrated Explainable Artificial Intelligence for Fetal Ultrasound Plane Classification: A Systematic Review

  • 提出CALIB-XFUS框架,整合校准、解释可信度与公平性评估
  • 78项研究汇总显示平均准确率达93%,但仅24%报告模型校准
  • 适合医疗AI开发者及监管机构参考,推动合规化落地

胎儿超声是产前护理的核心,准确识别一组标准解剖平面对于生物测量、生长监测和结构异常检测至关重要。深度学习分类器在受控基准上已达到或超过专家水平,但多数模型仍不透明且校准不足,导致临床医生缺乏可靠的置信度和真实解释,难以安全使用。本研究系统回顾了2015年1月1日至2026年4月30日间发表的78项研究,这些研究将自动化胎儿平面分类与可解释性或预测不确定性量化相结合,遵循PRISMA 2020标准。六种标准平面的总体平衡准确率为0.93(95% CI 0.91–0.95),但仅有19项研究(24%)报告了模型校准,14项(18%)报告了选择性预测。我们提出CALIB-XFUS,一个包含22个条目的报告框架,用于在受监管的胎儿超声人工智能中实现校准、解释忠实性和公平性。该框架涵盖六个领域:临床任务与用途;数据集来源与代表性;模型与训练流程;校准与选择性预测;解释忠实性与临床验证;上市后监测。我们认为,在FDA良好机器学习实践原则与欧盟《人工智能法案》高风险义务下,具备不确定性校准、可信解释和公平性审计的胎儿超声AI,现已在技术上可行且成为监管预期。

原文摘要 · Abstract (English)

Fetal ultrasound is the cornerstone of antenatal care, and accurate recognition of a small set of standard anatomical planes underpins biometry, growth surveillance, and detection of structural anomalies. Deep learning classifiers now match or exceed expert accuracy on curated benchmarks, but most remain opaque and miscalibrated, leaving clinicians without the calibrated confidence or faithful explanations needed for safe decision support. We systematically reviewed 78 studies published between January 1, 2015 and April 30, 2026 that paired automated fetal plane classification with explainability or predictive uncertainty quantification, following PRISMA 2020. Pooled balanced accuracy across six standard planes was 0.93 (95% CI 0.91 to 0.95), but only 19 studies (24%) reported calibration and 14 (18%) reported selective prediction. We propose CALIB-XFUS, a 22-item reporting framework that operationalises calibration, explanation faithfulness, and fairness for regulated fetal ultrasound artificial intelligence. The framework spans six domains: clinical task and indication for use; dataset provenance and representativeness; model and training pipeline; calibration and selective prediction; explanation faithfulness and clinician validation; and post-market surveillance. We argue that uncertainty-calibrated, faithfully explained, and fairness-audited fetal ultrasound AI is now both technically feasible and regulatorily expected under the FDA Good Machine Learning Practice principles and the EU AI Act high-risk obligations.

医学影像可解释AI不确定性量化联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。