arXiv:2411.17595cs.LGstat.AP2024-11被引 2

AI可预测临床试验结果,但不同模型各有优劣。

Can artificial intelligence predict clinical trial outcomes?

  • 用大模型和HINT模型分析试验结果预测性能。
  • GPT-4o整体表现最好,但难识别失败案例;HINT在负样本识别上更准。
  • 适合想提升试验预测准确率的研究者或药企决策者。

本研究评估了大型语言模型(LLMs)与HINT模型在预测临床试验结果上的表现,重点指标包括平衡准确率、马修斯相关系数(MCC)、召回率和特异性。结果显示,GPT-4o在所有LLMs中表现最优,但与其他模型(如GPT-3.5、GPT-4mini、Llama3)一样,在识别负面结果方面存在困难。相比之下,HINT在负样本识别上表现突出,且对招募挑战等外部因素具有鲁棒性,但在肿瘤学试验中表现不佳,该类试验占数据集主要部分。LLMs在早期试验及简单终点(如总生存期,OS)上表现较好,而HINT在多阶段试验中保持稳定,并在复杂终点(如客观缓解率,ORR)上优势明显。持续时间分析显示,中长期试验的模型性能更好,其中GPT-4o表现出稳定性,而HINT则具备更高特异性。研究强调了将LLMs(如GPT-4o、Llama3)与HINT结合使用的互补潜力,建议采用混合方法以发挥GPT-4o的预测力和HINT的高特异性,提升临床试验结果预测能力。

原文摘要 · Abstract (English)

This study evaluates the performance of large language models (LLMs) and the HINT model in predicting clinical trial outcomes, focusing on metrics including Balanced Accuracy, Matthews Correlation Coefficient (MCC), Recall, and Specificity. Results show that GPT-4o achieves superior overall performance among LLMs but, like its counterparts (GPT-3.5, GPT-4mini, Llama3), struggles with identifying negative outcomes. In contrast, HINT excels in negative sample recognition and demonstrates resilience to external factors (e.g., recruitment challenges) but underperforms in oncology trials, a major component of the dataset. LLMs exhibit strengths in early-phase trials and simpler endpoints like Overall Survival (OS), while HINT shows consistency across trial phases and excels in complex endpoints (e.g., Objective Response Rate). Trial duration analysis reveals improved model performance for medium- to long-term trials, with GPT-4o and HINT displaying stability and enhanced specificity, respectively. We underscore the complementary potential of LLMs (e.g., GPT-4o, Llama3) and HINT, advocating for hybrid approaches to leverage GPT-4o's predictive power and HINT's specificity in clinical trial outcome forecasting.

AI预测临床试验大模型医学应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。