arXiv:2506.18819cs.CLcs.AI2025-06

评测大模型总结真实世界证据研究的能力,发现Gemini 2.5表现最优。

RWESummary: A Framework and Test for Choosing Large Language Models to Summarize Real-World Evidence (RWE) Studies

  • 构建专用于真实世界证据总结的评估框架与测试集。
  • 基于13项真实研究,验证Gemini 2.5在摘要任务中综合表现最佳。
  • 适合医疗研究自动化、医学AI评估方向的研究者参考。

大型语言模型(LLMs)在通用摘要和医学研究辅助方面已有广泛评估,但尚未针对真实世界证据(RWE)研究结构化输出的摘要任务进行专门评测。本文提出RWESummary,作为MedHELM框架(Bedi, Cui, Fuentes, Unell et al., 2025)的补充,用于评估此类任务中的模型表现。该框架包含一个应用场景和三项评估,覆盖医学研究摘要中常见的主要错误类型,并基于Atropos Health的私有数据开发。此外,我们利用RWESummary对比了内部RWE摘要工具中不同LLMs的表现。截至发表时,在13项不同的RWE研究上,Gemini 2.5系列模型(Flash与Pro)整体表现最佳。建议将RWESummary作为真实世界证据研究摘要的新型基准模型评估基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been extensively evaluated for general summarization tasks as well as medical research assistance, but they have not been specifically evaluated for the task of summarizing real-world evidence (RWE) from structured output of RWE studies. We introduce RWESummary, a proposed addition to the MedHELM framework (Bedi, Cui, Fuentes, Unell et al., 2025) to enable benchmarking of LLMs for this task. RWESummary includes one scenario and three evaluations covering major types of errors observed in summarization of medical research studies and was developed using Atropos Health proprietary data. Additionally, we use RWESummary to compare the performance of different LLMs in our internal RWE summarization tool. At the time of publication, with 13 distinct RWE studies, we found the Gemini 2.5 models performed best overall (both Flash and Pro). We suggest RWESummary as a novel and useful foundation model benchmark for real-world evidence study summarization.

大模型评测真实世界证据医学摘要LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。