arXiv:2608.21079cs.LG2026-08

用大模型生成妊娠不良结局的因果假设,边试边调更准

Causal Modeling of Adverse Pregnancy Outcomes via Adaptive LLM Proposals

  • 让大模型当‘猜想生成器’,结合数据评分迭代优化
  • 在真实临床数据上复现全部专家认定因果关系,还发现新关联
  • 适合医学因果推断、临床研究者,尤其数据少时更有效

妊娠不良结局(APOs)如早产和妊娠期糖尿病对母婴有长期影响,但其成因仍不明确。由于数据稀缺和领域知识不全,纯数据驱动方法失效,大模型输出也常矛盾。本文提出一种神经符号框架,将大模型视为自适应提案分布,生成因果假设并基于实证数据打分;高分图用于更新模型上下文,引导后续生成聚焦更优假设空间。在真实临床数据集上评估该方法,与专家构建的因果图对比,成功恢复所有专家验证的边,并发现未被专家列出的新可能因果关系,或为靶向干预提供新洞见。

原文摘要 · Abstract (English)

Adverse Pregnancy Outcomes (APOs) such as preterm birth and gestational diabetes can have long-term consequences for both the mother and child, yet an understanding of their causes remains elusive. Causal discovery in this domain is especially challenging due to a paucity of data and incomplete domain knowledge. As a result, pure data-driven methods fail, and Large Language Model (LLM) outputs remain inconsistent or contradictory. We introduce a neurosymbolic framework for generating plausible causal hypotheses that iteratively combines the broad prior knowledge of LLMs with empirical scoring on data. Our method treats the LLM as an adaptive proposal distribution, generating hypotheses that are scored against empirical data; the resulting high-scoring graphs are then used to update the LLM's context, steering subsequent generations toward more promising regions of the hypothesis space. We evaluate our approach on a real-world clinical dataset for modeling APOs and their risk factors, comparing our results against an expert-constructed causal graph. Our method recovers all expert-validated edges and identifies additional plausible causal relations not previously listed by experts, potentially providing new insights for targeted interventions.

因果推断医学AI大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。