arXiv:2504.10397cs.AIcs.LG2025-04被引 3

用大模型生成医疗因果图,效果接近专家但有幻觉风险。

Can LLMs Assist Expert Elicitation for Probabilistic Causal Modeling?

  • 用大模型自动构建贝叶斯网络,替代人工专家推导因果关系。
  • 生成网络熵更低,预测更确定,优于专家和统计方法。
  • 适合医疗决策建模,但需警惕训练数据带来的偏差和幻觉。

本研究探讨大语言模型(LLMs)在生物识别与医疗应用中,作为提取结构化因果知识并辅助概率因果建模的替代方案。通过在医疗数据集上对比大模型生成的贝叶斯网络(BNs)与传统统计方法(如贝叶斯信息准则)的表现,采用结构方程模型(SEM)验证关系,并以熵、预测准确率及鲁棒性等指标评估网络结构。结果表明,大模型生成的贝叶斯网络熵值低于专家提取和统计生成的网络,表明其预测更具信心与精度。然而,上下文限制、幻觉依赖关系及训练数据潜在偏见等问题仍需深入研究。结论认为,大模型为概率因果建模中的专家征询开辟了新路径,有望提升决策透明度并降低不确定性。

原文摘要 · Abstract (English)

Objective: This study investigates the potential of Large Language Models (LLMs) as an alternative to human expert elicitation for extracting structured causal knowledge and facilitating causal modeling in biometric and healthcare applications. Material and Methods: LLM-generated causal structures, specifically Bayesian networks (BNs), were benchmarked against traditional statistical methods (e.g., Bayesian Information Criterion) using healthcare datasets. Validation techniques included structural equation modeling (SEM) to verifying relationships, and measures such as entropy, predictive accuracy, and robustness to compare network structures. Results and Discussion: LLM-generated BNs demonstrated lower entropy than expert-elicited and statistically generated BNs, suggesting higher confidence and precision in predictions. However, limitations such as contextual constraints, hallucinated dependencies, and potential biases inherited from training data require further investigation. Conclusion: LLMs represent a novel frontier in expert elicitation for probabilistic causal modeling, promising to improve transparency and reduce uncertainty in the decision-making using such models.

因果建模大模型应用医疗AI贝叶斯网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。