用大模型先验减少临床试验所需患者数,提升安全评估效率。
How many patients could we save with LLM priors?
- 用预训练大模型提取先验分布,直接指导贝叶斯建模
- 在真实临床数据上验证,预测性能优于传统方法
- 适合药物安全监测与高效临床试验设计人员
设想一个世界:借助大语言模型(LLM)编码的知识,临床试验可大幅减少所需患者数量以达到相同统计效力。我们提出一种用于多中心临床试验不良事件的分层贝叶斯建模新框架,利用LLM生成的先验分布。与生成合成数据的数据增强方法不同,本方法直接从模型获取参数化先验。通过系统的温度敏感性分析和在真实临床试验数据上的严格交叉验证,我们证明基于LLM的先验在预测性能上持续优于传统元分析方法。该方法为更高效、更依赖专家知识的临床试验设计铺平道路,显著降低实现稳健安全性评估所需的患者数量,有望改变药物安全监测与监管决策方式。
原文摘要 · Abstract (English)
Imagine a world where clinical trials need far fewer patients to achieve the same statistical power, thanks to the knowledge encoded in large language models (LLMs). We present a novel framework for hierarchical Bayesian modeling of adverse events in multi-center clinical trials, leveraging LLM-informed prior distributions. Unlike data augmentation approaches that generate synthetic data points, our methodology directly obtains parametric priors from the model. Our approach systematically elicits informative priors for hyperparameters in hierarchical Bayesian models using a pre-trained LLM, enabling the incorporation of external clinical expertise directly into Bayesian safety modeling. Through comprehensive temperature sensitivity analysis and rigorous cross-validation on real-world clinical trial data, we demonstrate that LLM-derived priors consistently improve predictive performance compared to traditional meta-analytical approaches. This methodology paves the way for more efficient and expert-informed clinical trial design, enabling substantial reductions in the number of patients required to achieve robust safety assessment and with the potential to transform drug safety monitoring and regulatory decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。