arXiv:2608.26109cs.AI2026-08

用智能体流程解释重症死亡预测,比单一大模型更可信。

Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

论文配图:Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset
图 1 · 摘自论文原文
  • 设计四步智能体流程,分步完成数据解读、指南验证与解释生成。
  • 在38个案例中,智能体流程零泄露,且指南契合度更高。
  • 适合临床风险解释场景,需搭配归因方法确保安全可用。

机器学习模型可精准预测重症监护室(ICU)死亡率,但仅靠特征归因方法难以提供临床所需叙事。大语言模型(LLM)或能弥补此差距,多步骤智能体流程因其可分离数据解读、指南核查与最终解释而具有可行性。本研究保留原始的独立模型与智能体对比设计,并更明确呈现主要临床发现。使用保留的eICU Demo数据集(2,353例ICU住院;8.1%死亡率),XGBoost模型达到AUROC 0.855(95%置信区间0.796–0.906)和AUPRC 0.332(95%置信区间0.217–0.494)。在38例分层解释子集上,独立LLM产生1例明确结果泄露,而四步智能体流程无一泄露。在14例与SHAP复核子集重叠的病例中,独立LLM的SHAP一致性(平均杰卡德指数0.171对0.077)和方向一致性(92.9%对78.6%)更高,但智能体流程在指南契合度(0.762对0.143)、价值具体性(0.236对0.143)和合理性(0.700对0.671)方面更优。临床提示:智能体分解可能提升安全相关契合度与个体化细节,但应在高风险解释前结合归因检查使用。

原文摘要 · Abstract (English)

Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation. This revised feasibility study preserves the original standalone-versus-agentic comparison while making the main clinical findings more explicit. Using the retained local eICU Demo artifact set (2,353 ICU stays; 8.1\% mortality), XGBoost achieved an AUROC of 0.855 (95\% CI 0.796--0.906) and an AUPRC of 0.332 (95\% CI 0.217--0.494). On a stratified 38-case explanation subset, the standalone LLM produced 1 explanation with explicit outcome leakage, whereas the four-step agentic pipeline produced none. Among the 14 cases that overlapped with the SHAP review subset, the standalone LLM showed higher SHAP alignment (mean Jaccard 0.171 versus 0.077) and higher direction consistency (92.9\% versus 78.6\%), while the agentic pipeline showed higher guideline grounding (0.762 versus 0.143), higher value specificity (0.236 versus 0.143), and slightly higher plausibility (0.700 versus 0.671). Clinically, the results suggest that agentic decomposition may improve safety-relevant grounding and patient-specific detail, but it should be paired with attribution-based checks before use in high-stakes risk explanation.

重症预测智能体可解释性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。