arXiv:2509.22565cs.CLcs.AI2025-09被引 3

用检索增强方法提升AI写患者消息的准确性,避免医疗错误。

Retrieval-Augmented Guardrails for AI-Drafted Patient-Portal Messages: Error Taxonomy Construction and Large-Scale Evaluation

  • 构建5大类59种临床错误编码体系,基于专家审核
  • 检索历史消息对提升错误识别率,F1达0.500
  • 适合医疗AI安全评估、医生辅助系统研发者

通过电子病历门户的异步医患沟通正成为医生工作负担来源,促使研究者探索大语言模型(LLMs)辅助撰写回复。然而,LLM输出可能存在临床错误、信息遗漏或语气不当等问题,因此需具备鲁棒性评估。本文贡献有三:(1) 提出一个基于临床实际的错误本体,包含5个领域和59个细粒度错误代码,经归纳编码与专家裁定建立;(2) 开发检索增强评估流程(RAEC),利用语义相似的历史消息-回复对提升判断质量;(3) 采用两阶段DSPy提示架构,实现可扩展、可解释且分层的错误检测。该方法在孤立评估与参考机构档案中历史对的基础上评估超过1,500条患者消息。引入检索上下文后,在临床完整性与流程适配性等领域的错误识别能力提升。对100条消息的人工验证显示,上下文增强标签的吻合度(一致性=50%)与性能(F1=0.500)均优于基线(一致性=33%,F1=0.256),支持将本RAEC流程作为患者消息生成的AI防护机制。

原文摘要 · Abstract (English)

Asynchronous patient-clinician messaging via EHR portals is a growing source of clinician workload, prompting interest in large language models (LLMs) to assist with draft responses. However, LLM outputs may contain clinical inaccuracies, omissions, or tone mismatches, making robust evaluation essential. Our contributions are threefold: (1) we introduce a clinically grounded error ontology comprising 5 domains and 59 granular error codes, developed through inductive coding and expert adjudication; (2) we develop a retrieval-augmented evaluation pipeline (RAEC) that leverages semantically similar historical message-response pairs to improve judgment quality; and (3) we provide a two-stage prompting architecture using DSPy to enable scalable, interpretable, and hierarchical error detection. Our approach assesses the quality of drafts both in isolation and with reference to similar past message-response pairs retrieved from institutional archives. Using a two-stage DSPy pipeline, we compared baseline and reference-enhanced evaluations on over 1,500 patient messages. Retrieval context improved error identification in domains such as clinical completeness and workflow appropriateness. Human validation on 100 messages demonstrated superior agreement (concordance = 50% vs. 33%) and performance (F1 = 0.500 vs. 0.256) of context-enhanced labels vs. baseline, supporting the use of our RAEC pipeline as AI guardrails for patient messaging.

医疗AILLM评估检索增强错误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。