FDA批准的医疗AI设备报告中,可信AI证据严重缺失。
Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices

- 分析519份FDA报告,仅24.7%提及任何可信AI原则。
- 鲁棒性报告最常见(57.6%),可追溯性与可解释性不足(8.3%、3.5%)。
- 监管审批不能代表可信度,需建立标准化报告规范。
背景:人工智能/机器学习(AI/ML)赋能的医疗设备在医疗领域日益普及,其监管框架也在不断演变。随着这些系统深度融入临床决策,公众、患者和医生对其可信性要求越来越高。然而,公开的监管文件是否足以独立评估已获准AI系统的可信性尚不明确。方法:我们分析了2021至2025年间发布的1,105份FDA AI/ML医疗设备摘要报告,经关键词筛选后进行多阶段人工共识评审,识别六项FUTURE-AI原则(公平性、普适性、可追溯性、可用性、鲁棒性、可解释性)的文档证据。进行了描述性、时间趋势及临床领域分析。通过多变量逻辑回归评估获批年份或临床领域是否预测更高透明度(报告三个及以上原则)。结果:共纳入519份报告。可信AI报告覆盖率低且分布不均:近四分之一(24.7%)未提及任何原则,无一报告覆盖全部六项。鲁棒性报告率最高(57.6%),可追溯性(8.3%)与可解释性(3.5%)为显著短板。获批年份(OR 1.02,95% CI 0.88–1.19)与临床领域(OR 0.73,95% CI 0.46–1.15)均无法预测更高透明度。结论:FDA文件中存在广泛且持续的可信AI报告空白。仅凭监管批准不应视为可信性的代名词。必须建立贯穿AI全生命周期的标准化、可审计的报告机制,以支持独立评估与负责任的医疗AI应用。
原文摘要 · Abstract (English)
Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dimensions of trustworthy AI to support clinician, patient, and public trust. Whether publicly available regulatory documentation provides sufficient evidence to independently assess the trustworthiness of cleared AI systems remains unclear. Methods: We analysed FDA AI/ML-enabled medical device summary reports published between 2021 and 2025. Reports underwent automated keyword screening followed by multi-stage manual consensus review to identify documented evidence for the six FUTURE-AI principles: Fairness, Universality, Traceability, Usability, Robustness, and Explainability. Descriptive, temporal, and clinical-domain analyses were performed. Multivariable logistic regression assessed whether year of clearance or clinical domain predicted higher reporting transparency, defined as evidence reported for three or more principles. Results: Of 1,105 FDA summary reports screened, 519 were included. Trustworthy AI reporting was limited and uneven. Nearly one quarter (24.7%) provided no evidence for any principle, and none documented evidence across all six. Robustness was most frequently reported (57.6%), while Traceability (8.3%) and Explainability (3.5%) were the most pronounced gaps. Neither year of clearance (OR 1.02, 95% CI 0.88-1.19) nor clinical domain (OR 0.73, 95% CI 0.46-1.15) predicted higher reporting transparency. Interpretation: Substantial, persistent trustworthy AI reporting gaps exist in FDA documentation. Regulatory approval alone should not be considered a proxy for trustworthiness. Standardised, audit-ready reporting across the AI lifecycle is needed to support independent assessment and responsible adoption of healthcare AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。