arXiv:2509.16372cs.AI2025-09被引 5

测试大模型在临床化验解读中的因果推理能力,发现GPT-o1表现更优。

Evaluation of Causal Reasoning for Large Language Models in Contextualized Clinical Scenarios of Laboratory Test Interpretation

  • 用珍珠因果阶梯框架评估大模型对化验结果的因果推理
  • GPT-o1在干预与反事实推理上准确率超Llama-3.2-8b-instruct,AUROC达0.84
  • 适合医疗AI评估者参考,尤其关注临床决策中的因果逻辑

本研究基于珍珠因果阶梯(关联、干预、反事实)框架,评估大语言模型在99个临床化验场景中的因果推理能力。测试了血红蛋白A1c、肌酐、维生素D等常见检验项目,关联年龄、性别、肥胖、吸烟等因果因素。对比GPT-o1与Llama-3.2-8b-instruct两个模型,由四位医学专家评估。GPT-o1整体判别性能更优(AUROC=0.80±0.12),高于Llama-3.2-8b-instruct(0.73±0.15);在关联(0.75 vs 0.72)、干预(0.84 vs 0.70)和反事实推理(0.84 vs 0.69)任务中均领先。敏感度(0.90 vs 0.84)与特异度(0.93 vs 0.80)也更高。两者在干预类问题表现最佳,反事实类最差,尤其在结果改变场景中。结果表明GPT-o1具备更一致的因果推理能力,但仍需优化方可用于高风险临床场景。

原文摘要 · Abstract (English)

This study evaluates causal reasoning in large language models (LLMs) using 99 clinically grounded laboratory test scenarios aligned with Pearl's Ladder of Causation: association, intervention, and counterfactual reasoning. We examined common laboratory tests such as hemoglobin A1c, creatinine, and vitamin D, and paired them with relevant causal factors including age, gender, obesity, and smoking. Two LLMs - GPT-o1 and Llama-3.2-8b-instruct - were tested, with responses evaluated by four medically trained human experts. GPT-o1 demonstrated stronger discriminative performance (AUROC overall = 0.80 +/- 0.12) compared to Llama-3.2-8b-instruct (0.73 +/- 0.15), with higher scores across association (0.75 vs 0.72), intervention (0.84 vs 0.70), and counterfactual reasoning (0.84 vs 0.69). Sensitivity (0.90 vs 0.84) and specificity (0.93 vs 0.80) were also greater for GPT-o1, with reasoning ratings showing similar trends. Both models performed best on intervention questions and worst on counterfactuals, particularly in altered outcome scenarios. These findings suggest GPT-o1 provides more consistent causal reasoning, but refinement is required before adoption in high-stakes clinical applications.

大模型因果推理临床决策医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。