构建双流记忆架构,实时比对患者自述与病历数据,提升健康助手安全性。
Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

- 分离患者叙述与电子病历,通过专用引擎检测偏差
- 在675次会话中识别出84.4%的设计偏差,关键错误召回率达86.7%
- 揭示记忆提取过程是错误传递主因,适合医疗AI安全研究者
随着大语言模型代理从单次交互转向长期健康管理,其记忆系统面临核心挑战:需协调两种不完美的信息源。患者自述虽实时但易受回忆偏差影响,而电子健康记录(EHR)虽医学可信却常滞后。通用代理记忆系统为追求连贯性,常以最新用户陈述覆盖旧信息,这在临床场景中可能导致安全风险。本文提出双流记忆架构,严格分离患者叙事与结构化临床记录(FHIR),由专门的校正引擎评估每条提取记忆,按类型、严重程度及涉及的具体FHIR资源分类偏差。在26名患者、675次纵向健康辅导会话上测试,采用真实医患对话与合成的FHIR基准临床情景混合数据集。孤立测试中,该引擎可检测84.4%的设计临床偏差,关键安全错误召回率达86.7%。通过在同一数据上联合评估提取与校正,直接量化出13.6%的错误级联,溯源显示问题主要源于非结构化对话中临床细节在记忆提取阶段丢失,而非下游分类错误。研究证实,在临床部署中验证患者自述与病历的一致性,既可行也必要。
原文摘要 · Abstract (English)
As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memory architectures face a critical challenge: reconciling two imperfect sources of truth. The patient's evolving self-report is current but prone to recall bias, while the Electronic Health Record (EHR) is medically validated but frequently stale. General-purpose agent memory systems optimize for coherence by overwriting older facts with the user's latest statement, a pattern that risks safety failures when applied to clinical data. We introduce a Dual-Stream Memory Architecture that strictly separates the patient narrative from the structured clinical record (FHIR), governed by a dedicated Reconciliation Engine that evaluates every extracted memory against the patient's FHIR profile and classifies discrepancies by type, severity, and the specific FHIR resources involved. We evaluate this architecture on 26 patients across 675 longitudinal wellness coaching sessions, using a hybrid dataset that interleaves real provider-patient transcripts with synthetic, FHIR-grounded clinical scenarios. In isolated testing, the engine detects 84.4% of designed clinical discrepancies with 86.7% safety-critical recall. By coupling extraction and reconciliation evaluation on the same data, we directly quantify a 13.6% error cascade, tracing the degradation to clinical details lost during memory extraction from unstructured conversation rather than to downstream classification errors. These findings establish that validating patient-reported memories against clinical records is both feasible and necessary for safe deployment of longitudinal health agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。