AI助手PowerGPT让临床试验样本量计算更准更快
Empowering Clinical Trial Design through AI: A Randomized Evaluation of PowerGPT
- 用大模型+统计引擎自动选检验方法和算样本量
- 准确率94.1%、完成率99.3%,比人工快一倍以上
- 适合医生、研究人员等非统计专家使用
临床研究中的功效分析需要精确的样本量计算,但其复杂性和对统计学专业知识的依赖阻碍了众多研究者。我们提出PowerGPT,一个集成大型语言模型(LLMs)与统计引擎的AI系统,用于自动化临床试验设计中的检验方法选择和样本量估算。在一项随机对照试验中评估其效果发现,PowerGPT显著提升了任务完成率(检验选择:99.3% vs. 88.9%,样本量计算:99.3% vs. 77.8%),提高了准确性(样本量估算:94.1% vs. 55.4%,p < 0.001),并缩短了平均耗时(4.0分钟 vs. 9.3分钟,p < 0.001)。这些优势在不同统计检验中保持一致,且对统计学家与非统计学家均有效,有助于弥合专业差距。目前该系统已在多个机构部署,为临床研究提供了可扩展的、智能化的统计功效分析方案。
原文摘要 · Abstract (English)
Sample size calculations for power analysis are critical for clinical research and trial design, yet their complexity and reliance on statistical expertise create barriers for many researchers. We introduce PowerGPT, an AI-powered system integrating large language models (LLMs) with statistical engines to automate test selection and sample size estimation in trial design. In a randomized trial to evaluate its effectiveness, PowerGPT significantly improved task completion rates (99.3% vs. 88.9% for test selection, 99.3% vs. 77.8% for sample size calculation) and accuracy (94.1% vs. 55.4% in sample size estimation, p < 0.001), while reducing average completion time (4.0 vs. 9.3 minutes, p < 0.001). These gains were consistent across various statistical tests and benefited both statisticians and non-statisticians as well as bridging expertise gaps. Already under deployment across multiple institutions, PowerGPT represents a scalable AI-driven approach that enhances accessibility, efficiency, and accuracy in statistical power analysis for clinical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。