首个针对生成式AI临床摘要的安全风险评估框架,可系统识别潜在医疗隐患。
Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content
- 基于FMECA构建风险分类体系,量化故障发生、严重性与可发现性。
- 在36份真实出院小结上验证,专家评分一致性达中等至良好水平。
- 适合医疗AI安全评审、临床决策支持系统开发者使用。
大型语言模型(LLM)在临床文本摘要中的应用日益广泛,但系统化的患者安全风险评估方法仍不足。本文首次将故障模式、影响与关键性分析(FMECA)框架应用于生成式临床内容,开发并验证了一个新型FMECA评估工具。研究由8名跨学科专家通过文献回顾与头脑风暴建立故障模式分类体系,并将发生率、严重性和可检测性转化为五级量表。该框架被用于评估4位患者共36份由GPT-OSS 120B生成的出院小结,数据来自日内瓦大学医院真实临床记录。两名评审员分两轮独立标注,评估组间一致性及可用性。结果显示:最终框架包含14类故障模式;专家间一致性在第二轮提升至中等到较高水平,严重性与可检测性评分达成良好一致性;系统可用性得分79.2/100,评价者信心高。本研究提出首个系统化评估生成式临床摘要患者安全风险的FMECA框架,为医疗AI风险管控提供可复现的方法支撑。
原文摘要 · Abstract (English)
Objectives: Large language models (LLMs) are increasingly used for clinical text summarization, yet structured methods to assess associated patient safety risks remain limited. Failure Mode, Effects, and Criticality Analysis (FMECA) provides a proactive framework for systematic risk identification but has not been adapted to LLM-generated clinical content. This study aimed to develop and validate a novel FMECA framework for the prospective assessment of patient safety risks in LLM-generated clinical summaries. Materials and Methods: An interdisciplinary expert panel (n = 8) developed a taxonomy of failure modes through literature review and brainstorming. Standard FMECA dimensions (occurrence, severity, detectability) were adapted into 5-point ordinal scales. The framework was applied to 36 discharge summaries from four patients, generated by an open LLM (GPT-OSS 120B) using real-world clinical data from the Geneva University Hospitals. Reviewers independently annotated the summaries across two rounds. Inter-rater reliability was assessed at failure mode, severity and detectability score levels. Usability and content validity were evaluated using an adapted System Usability Scale and structured feedback. Results: The final framework comprised 14 failure modes organized into categories. Inter-rater agreement improved between rounds, reaching moderate-to-substantial agreement for failure mode identification and good agreement for severity and detectability scoring. Usability was rated as good (mean SUS: 79.2/100), with high evaluator confidence. Discussion and Conclusion: This study presents the first FMECA-based framework for systematic patient safety risk assessment of LLM-generated clinical summaries. The framework provides a structured and reproducible method for identifying clinically relevant risks caused by these summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。