用可追溯的LLM框架帮医疗设备安全团队高效整合证据,减少专家负担。
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices

- 构建证据锚定的LLM系统,关联设备文档与知识库
- 支持生成候选安全项并自动检查错误与重复,准确率提升显著
- 适合医疗设备合规团队、监管机构及需要可追溯性的研发人员
医疗设备日益软件化、联网化和智能化,其开发需符合ISO 14971风险管理和IEC 62304软件标准。相关证据必须在需求、设计、变更、验证结果、投诉及上市后数据间保持一致,但此类工作成本高且依赖稀缺的安全与领域专家。大语言模型(LLM)因安全工作高度依赖文档,或可降低部分负担。然而,现有研究多聚焦孤立方法,依赖通用提示或公开示例,缺乏对来源链接、可追溯性、不确定性处理、生命周期更新及专家评审记录的支持,难以用于受监管的医疗设备开发。本文认为核心问题并非生成安全文本,而是提供有依据的安全知识支持。为此提出一种证据锚定框架:连接设备产物、可控知识存储与检索、特定方法生成候选安全项、批判性与不确定性检查,并记录专家评审。该框架为专家决策准备、链接、校验和更新候选安全文档,不判定设备安全性,也不提供监管批准。还提出使用非公开或新构建的医疗设备案例研究与专家参考分析,评估覆盖度、正确性、相关性、可追溯性、重复率、无依据主张及评审工作量。
原文摘要 · Abstract (English)
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety work is highly document-based. However, current LLM-based safety-engineering studies often address isolated methods, rely on generic prompting or public examples, and provide limited support for source links, traceability, uncertainty handling, lifecycle updates, and recorded expert review. This limits their use in regulated medical-device development. This paper argues that the central research problem is not safety-text generation, but source-linked safety-knowledge support. We propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making. It does not decide whether a device is safe and does not provide regulatory approval. We also outline an evaluation strategy using non-public or newly built medical-device case studies and expert reference analyses to assess coverage, correctness, relevance, traceability, duplicate rate, unsupported claims, and review effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。