arXiv:2607.05032cs.CL2026-07

用智能大模型自动评估糖尿病患者病情严重程度,效果媲美专有系统。

Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

  • 构建两阶段智能体框架,整合多源临床证据进行推理
  • 在4886例合成数据上达到中等以上一致性,预测死亡率与并发症显著有效
  • 开源版本表现接近闭源系统,适合医疗研究与个性化诊疗应用

疾病严重程度是多维度复杂概念,传统规则难以在电子健康记录(EHR)中准确捕捉。MOSAIC是一种两阶段的代理型大语言模型(LLM)框架,以2型糖尿病(T2D)为验证案例。在合成队列SyntheticMass(开放权重N=4,886;封闭权重N=200)上评估,对比三种算法基准(DCSI、DiSSCo、Cooper),以及全因死亡率和新发并发症。该框架覆盖了原有方法未包含的领域,如基于生物标志物的血糖分期、β细胞功能及健康的社会决定因素。开放权重版本与专有管道表现相当(加权克朗巴系数0.773),与Cooper(κ=0.597)、DCSI(κ=0.534)达中等一致,与DiSSCo(κ=0.320)达良好一致。代理型(类型1)层级在全因死亡率上显著分离(log-rank p < 0.001;粗略风险比1.6–2.4),上层出现非单调分离,且新发并发症呈反向梯度,符合易感人群耗竭现象。与固定规则执行相比(MOSAIC Frozen;κ=0.428),表明其具备超越规则的推理能力。结论:MOSAIC证明代理型LLM可从结构化EHR生成并应用临床有意义的严重程度表型。扩展至其他具有多维严重程度特征的疾病值得进一步研究。

原文摘要 · Abstract (English)

Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Agentic large language model (LLM) systems could synthesise clinical evidence and reason over EHRs, but remain unevaluated for this task. Methods: MOSAIC is a two-phase agentic LLM framework for severity phenotyping, using type 2 diabetes (T2D) as a proof-of-concept. MOSAIC was evaluated on a synthetic cohort (SyntheticMass; open-weight N = 4,886; closed-weight N = 200) against three algorithmic ground truths (DCSI, DiSSCo, Cooper) and against all-cause mortality and incident complications. Open-weight (locally deployable) and proprietary pipelines were also compared. Results: The generated framework spanned domains absent from the comparators, including biomarker-based glycaemic staging, beta-cell function, and social determinants of health. Open-weight MOSAIC matched the proprietary pipeline (closed- vs open-weight weighted kappa = 0.773) and reached moderate agreement with Cooper (kappa = 0.597) and DCSI (kappa = 0.534) and fair agreement with DiSSCo (kappa = 0.320). Agent-based (Type 1) tiers showed significant separation of all-cause mortality (log-rank p < 0.001; crude hazard ratios 1.6-2.4 for non-Baseline tiers), with non-monotonic separation at the upper tiers, and an inverse gradient for incident complications (log-rank p < 0.001) consistent with depletion of susceptibles. Agentic classification also diverged from deterministic execution of the same rubric (MOSAIC Frozen; kappa = 0.428), indicating reasoning beyond fixed rules. Conclusion: MOSAIC shows agentic LLM systems can generate and apply clinically meaningful severity phenotypes from structured EHR data in T2D. Extending it to other diseases with similarly multidimensional severity warrants further research.

医疗AI大模型严重程度评估糖尿病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。