用双模块框架提升病患出院计划的准确与可靠
Planner-Auditor Twin: Agentic Discharge Planning with FHIR-Based LLM Planning, Guideline Recall, Optional Caching and Self-Improvement
- 拆分生成与校验:大模型拟计划,规则模块逐项审核
- 覆盖率达86%,高置信度遗漏减少,校准误差显著下降
- 适合医疗AI研发者,尤其关注安全与可验证性的临床应用
大型语言模型在临床出院规划中展现出潜力,但受限于幻觉、遗漏和置信度失准。本文提出一种自改进、可选缓存的规划-审计器框架,通过解耦生成与确定性验证,提升安全性与可靠性。基于MIMIC-IV-on-FHIR构建代理式回溯评估流程,每个患者由规划器(LLM)生成带置信度估计的结构化出院计划,审计器为确定性模块,评估多任务覆盖率,追踪校准性(Brier得分、ECE代理指标),并监控动作分布漂移。支持两阶段自我改进:(i) 启用时的单轮再生,(ii) 高置信度低覆盖率案例的跨轮次差异缓冲与重播。结果表明,尽管上下文缓存提升性能,但自我改进环是主要驱动力,使任务覆盖率从32%提升至86%;校准显著改善,Brier/ECE降低,高置信度遗漏减少。差异缓冲进一步通过重播修正持续存在的高置信度遗漏。讨论指出,反馈驱动再生与目标重播有效抑制遗漏,提升置信度可靠性。将LLM规划器与基于规则的观测审计器分离,实现系统性可靠性度量与无需重训练的安全迭代。结论:该框架通过可互操作的FHIR数据访问与确定性审计,为更安全的自动化出院规划提供可行路径,具备可复现的消融分析与可靠性导向评估。
原文摘要 · Abstract (English)
Objective: Large language models (LLMs) show promise for clinical discharge planning, but their use is constrained by hallucination, omissions, and miscalibrated confidence. We introduce a self-improving, cache-optional Planner-Auditor framework that improves safety and reliability by decoupling generation from deterministic validation and targeted replay. Materials and Methods: We implemented an agentic, retrospective, FHIR-native evaluation pipeline using MIMIC-IV-on-FHIR. For each patient, the Planner (LLM) generates a structured discharge action plan with an explicit confidence estimate. The Auditor is a deterministic module that evaluates multi-task coverage, tracks calibration (Brier score, ECE proxies), and monitors action-distribution drift. The framework supports two-tier self-improvement: (i) within-episode regeneration when enabled, and (ii) cross-episode discrepancy buffering with replay for high-confidence, low-coverage cases. Results: While context caching improved performance over baseline, the self-improvement loop was the primary driver of gains, increasing task coverage from 32% to 86%. Calibration improved substantially, with reduced Brier/ECE and fewer high-confidence misses. Discrepancy buffering further corrected persistent high-confidence omissions during replay. Discussion: Feedback-driven regeneration and targeted replay act as effective control mechanisms to reduce omissions and improve confidence reliability in structured clinical planning. Separating an LLM Planner from a rule-based, observational Auditor enables systematic reliability measurement and safer iteration without model retraining. Conclusion: The Planner-Auditor framework offers a practical pathway toward safer automated discharge planning using interoperable FHIR data access and deterministic auditing, supported by reproducible ablations and reliability-focused evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。