用多智能体大模型自动生成临床级特征,省时48倍且效果媲美专家
Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models
- 设计多智能体系统模仿医生逐轮审阅流程,自动从病历中提取特征
- 在前列腺癌预测中AUC达0.767,接近人工标注的0.762,优于多种基线方法
- 无需调参即可跨病种扩展,适合需高效构建临床预测模型的研究者
构建精准临床预测模型常受限于从非结构化电子病历(EHR)笔记中提取有意义结构化特征的难度。本研究首先建立患者级别的临床特征生成(CFG)协议,由领域专家手动审查147例前列腺癌患者病历,提取精细特征作为高保真基准。基于此,提出透明的多智能体大语言模型系统SNOW(Scalable Note-to-Outcome Workflow),可自主模拟专家的迭代推理与验证流程。在5年癌症复发预测中,SNOW(AUC-ROC 0.767)性能接近人工CFG(0.762),优于结构化基线、医生引导的LLM提取及六种表示特征生成(RFG)方法。配置后,SNOW仅用12小时完成全患者特征表生成,结合5小时医生监督,相较人工流程减少约48倍人力。为测试可扩展性,将SNOW部署于外部心衰(射血分数保留型,HFpEF)队列(MIMIC-IV,n=2,084),未进行任务特化调优即生成预后特征,在30天(SNOW: 0.851)和1年(SNOW: 0.763)死亡预测中表现超越基线与RFG方法。结果表明,模块化多智能体系统可规模化生成专家级临床特征,实现对非结构化病历文本的可解释利用,并保持跨场景泛化能力。
原文摘要 · Abstract (English)
Developing accurate clinical prediction models is often bottlenecked by the difficulty of deriving meaningful structured features from unstructured EHR notes, a process that traditionally requires manual, unscalable clinical abstraction. In this study, we first established a rigorous patient-level Clinician Feature Generation (CFG) protocol, in which domain experts manually reviewed notes to define and extract nuanced features for a cohort of 147 patients with prostate cancer. As a high-fidelity ground truth, this labor-intensive process provided the blueprint for SNOW (Scalable Note-to-Outcome Workflow), a transparent multi-agent large language model (LLM) system designed to autonomously mimic the iterative reasoning and validation workflow of clinical experts. On 5-year cancer recurrence prediction, SNOW (AUC-ROC 0.767) achieved performance comparable to manual CFG (0.762) and outperformed structured baselines, clinician-guided LLM extraction, and six representational feature generation (RFG) approaches. Once configured, SNOW produced the full patient-level feature table in 12 hours with 5 hours of clinician oversight, reducing human expert effort by approximately 48-fold versus manual CFG. To test scalability where manual CFG is infeasible, we deployed SNOW on an external heart failure with preserved ejection fraction (HFpEF) cohort from MIMIC-IV (n=2,084); without task-specific tuning, SNOW generated prognostic features that outperformed baseline and RFG methods for 30-day (SNOW: 0.851) and 1-year (SNOW: 0.763) mortality prediction. These results demonstrate that a modular LLM agent-based system can scale expert-level feature generation from clinical notes, while enabling interpretable use of unstructured EHR text in outcome prediction and preserving generalizability across a variety of settings and conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。