arXiv:2505.16097cs.AI2025-05被引 2

用百万临床试验数据训练大模型,提升医学研究智能化水平

Developing Large Language Models for Clinical Research Using One Million Clinical Trials

  • 基于160万临床试验数据构建结构化资源TrialPanorama,融合医学本体与文献
  • 自研80亿参数模型在8项临床任务中超越700亿通用模型,提升超20%
  • 适合医疗AI研究者、临床试验设计人员快速提升决策效率

构建人工智能用于临床研究需要全面的数据基础以支持模型训练与严格评估。本文提出TrialPanorama,一个大规模结构化资源,整合了来自十五个全球注册库的160万条临床试验记录,并关联生物医学本体及相关文献。为验证其价值,我们构建了涵盖8项关键临床研究任务的流水线,生成15.2万组训练与测试样本。其中3项支持系统综述流程(研究检索、筛选、证据总结),5项聚焦试验设计优化(臂设计、纳入标准制定、终点选择、样本量估算、完成度评估与合理性分析)。基准测试表明,通用大语言模型在临床推理能力上表现有限;而我们在TrialPanorama上通过监督微调与强化学习训练的80亿参数模型,在所有8项任务中均优于700亿参数通用模型,相对提升分别为73.7%、67.6%、38.4%、37.8%、26.5%、20.7%、20.0%和5.2%。我们认为TrialPanorama为未来临床研究AI的规模化发展提供了坚实基础。

原文摘要 · Abstract (English)

Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation that supports model training and rigorous evaluation. Here, we introduce TrialPanorama, a large-scale structured resource that aggregates 1.6M clinical trial records from fifteen global registries and links them with biomedical ontologies and associated literature. To demonstrate its utility, we build a pipeline that constructs 152K training and testing samples for eight key clinical research tasks. Three tasks support systematic review workflows, including study search, study screening, and evidence summarization. Five tasks focus on trial design and optimization, including arm design, eligibility criteria design, endpoint selection, sample size estimation, and trial completion assessment and rationalization. Benchmarking cutting-edge large language models (LLMs) reveals that generic LLMs have limited capability in clinical reasoning. In contrast, an 8B LLM we developed on TrialPanorama using supervised finetuning and reinforcement learning wins over the 70B generic counterparts in all eight tasks, with a relative improvement of 73.7%, 67.6%, 38.4%, 37.8%, 26.5%, 20.7%, 20.0%, 18.1%, and 5.2%, respectively. We envision that TrialPanorama provides a solid foundation for future scaling of AI for clinical research.

临床研究大模型知识图谱医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。