arXiv:2608.10339stat.MEcs.AI2026-08

用大模型辅助专家判断,精准估算医院干预措施省下的住院时间。

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

论文配图:Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement
图 1 · 摘自论文原文
  • 结合甘特图与因果图,仅对数据无法识别部分请求专家输入。
  • 在11项干预措施中,大模型生成结果与专家一致率高。
  • 适合缺乏历史数据、需临床推理的复杂干预效果评估。

医院质量改进常面临多个候选干预措施,但现有方法难以估计和排序其因果效应。本文聚焦平均住院时长(LOS)这一核心指标及其因果效应——平均节省时间。传统定性方法依赖专家判断患者路径,易受认知偏差影响;定量方法依赖数据模型,在干预无历史数据或机制复杂时失效。为此提出专家引导的g-计算(egg-computation),将甘特图与因果DAG结合,通过变体g-计算仅在数据无法识别处请求专家输入。为实现可扩展性,开发了基于大模型的辅助流程,可靠放大专家推理能力。模拟实验表明,在患者因果结构多样时,egg-computation优于传统方法;在一家城市安全网医院对11项候选干预的实证研究中,大模型生成的路径图与节省时间估计与人类专家高度一致。该框架可推广至其他需甘特图建模因果机制的场景。

原文摘要 · Abstract (English)

Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of the most standard hospital metrics, the average length of stay (LOS), and its causal estimand, the average time saved. To characterize this causal effect, qualitative approaches rely on expert judgment to map patient trajectories, making them susceptible to cognitive biases; quantitative approaches rely on data-driven models, which fail when interventions are hypothetical with no historical data or have complex causal mechanisms that require clinical reasoning rather than data alone. We propose expert-guided g-computation, or egg-computation, which combines the complementary strengths of both approaches by connecting the Gantt charts commonly used to map patient trajectories with the causal DAG literature. We introduce a causal model over Gantt charts and establish identification using a variant of g-computation that seeks expert input only for components unidentifiable from data. To make egg-computation practical, we develop an LLM-assisted pipeline that reliably scales up expert reasoning. In simulations, egg-computation outperforms conventional causal inference methods when patients have diverse causal structures and intervention mechanisms. In a study of eleven candidate QI interventions at an urban safety-net hospital, the LLM pipeline generated graphs and time-saving estimates highly concordant with those of human experts. Beyond healthcare, egg-computation is a broadly applicable framework for estimating the average time saved for candidate interventions whose causal mechanisms can be represented using Gantt charts.

因果推断医疗优化大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。