用多智能体协作模拟医生会诊,提升病历问题识别准确率。
Automated Clinical Problem Detection from SOAP Notes using a Collaborative Multi-Agent LLM Architecture
- 构建动态协作的多智能体系统,模拟临床团队诊断过程。
- 在420份MIMIC-III病历上,对心衰、肾损伤和败血症识别效果更优。
- 适合需要高可靠性与可解释性的临床决策支持场景。
准确解读临床叙事对患者诊疗至关重要,但病历内容复杂,自动化面临挑战。尽管大语言模型(LLMs)展现出潜力,单一模型方法在高风险临床任务中仍缺乏鲁棒性。本文提出一种协作式多智能体系统(MAS),模拟临床会诊团队,仅分析SOAP病历中的主观(S)和客观(O)部分,实现从原始数据到诊断判断的推理。由管理代理动态调度专科代理,通过层级化、迭代式辩论达成共识。在包含420例MIMIC-III病历的定制数据集上,该系统在识别充血性心力衰竭、急性肾损伤和败血症方面表现优于单智能体基线。定性分析显示,该结构能有效暴露并权衡矛盾证据,但偶发群体思维现象。通过模拟临床团队推理,本系统为更准确、鲁棒且可解释的临床决策支持工具提供了新路径。
原文摘要 · Abstract (English)
Accurate interpretation of clinical narratives is critical for patient care, but the complexity of these notes makes automation challenging. While Large Language Models (LLMs) show promise, single-model approaches can lack the robustness required for high-stakes clinical tasks. We introduce a collaborative multi-agent system (MAS) that models a clinical consultation team to address this gap. The system is tasked with identifying clinical problems by analyzing only the Subjective (S) and Objective (O) sections of SOAP notes, simulating the diagnostic reasoning process of synthesizing raw data into an assessment. A Manager agent orchestrates a dynamically assigned team of specialist agents who engage in a hierarchical, iterative debate to reach a consensus. We evaluated our MAS against a single-agent baseline on a curated dataset of 420 MIMIC-III notes. The dynamic multi-agent configuration demonstrated consistently improved performance in identifying congestive heart failure, acute kidney injury, and sepsis. Qualitative analysis of the agent debates reveals that this structure effectively surfaces and weighs conflicting evidence, though it can occasionally be susceptible to groupthink. By modeling a clinical team's reasoning process, our system offers a promising path toward more accurate, robust, and interpretable clinical decision support tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。