评测大模型结合因果特征对临床预测的提升效果
REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks
- 构建新基准REACT-LLM,评估大模型与因果特征融合性能
- 15个大模型+6个传统模型+3种因果发现算法对比,7个临床结局
- 因果特征融入大模型提升有限,但揭示了未来协同潜力
大型语言模型(LLMs)和因果学习在临床决策中均具潜力,但二者结合的效果尚不明确,主要因缺乏系统性基准评估其在临床风险预测中的整合。真实医疗场景中,识别对结果具有因果影响的特征对生成可行动、可信的预测至关重要。尽管近期研究显示大模型具备初步因果推理能力,但缺乏全面基准来评估其在临床风险预测中利用因果特征的表现。为此,我们提出REACT-LLM基准,旨在评估将大模型与因果特征结合是否能提升临床预后性能,并可能超越传统机器学习方法。该基准在两个真实世界数据集上评估7个临床结局,对比15个主流大模型、6个传统机器学习模型和3种因果发现(CD)算法。结果显示,虽然大模型在临床预后任务中表现合理,但尚未超越传统机器学习方法。将因果发现算法提取的因果特征融入大模型带来的性能提升有限,主要源于多数因果发现方法依赖严格假设,在复杂临床数据中常被违反。尽管直接融合效果有限,该基准揭示了更具前景的协同路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and causal learning each hold strong potential for clinical decision making (CDM). However, their synergy remains poorly understood, largely due to the lack of systematic benchmarks evaluating their integration in clinical risk prediction. In real-world healthcare, identifying features with causal influence on outcomes is crucial for actionable and trustworthy predictions. While recent work highlights LLMs' emerging causal reasoning abilities, there lacks comprehensive benchmarks to assess their causal learning and performance informed by causal features in clinical risk prediction. To address this, we introduce REACT-LLM, a benchmark designed to evaluate whether combining LLMs with causal features can enhance clinical prognostic performance and potentially outperform traditional machine learning (ML) methods. Unlike existing LLM-clinical benchmarks that often focus on a limited set of outcomes, REACT-LLM evaluates 7 clinical outcomes across 2 real-world datasets, comparing 15 prominent LLMs, 6 traditional ML models, and 3 causal discovery (CD) algorithms. Our findings indicate that while LLMs perform reasonably in clinical prognostics, they have not yet outperformed traditional ML models. Integrating causal features derived from CD algorithms into LLMs offers limited performance gains, primarily due to the strict assumptions of many CD methods, which are often violated in complex clinical data. While the direct integration yields limited improvement, our benchmark reveals a more promising synergy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。