arXiv:2603.22767cs.AIcs.CL2026-03被引 2

评测大模型在真实医疗数据中生成完整医学证据的能力

Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases

  • 构建基于MIMIC-IV的基准测试,要求模型生成树状结构的证据链
  • 六种模型在162项任务中最高成功率仅39.9%,开放模型为30.4%
  • 揭示了模型在全流程执行中的系统性缺陷,适合医学AI评估研究者

观察性研究可大规模生成临床可用证据,但在真实世界数据库中执行需贯穿队列构建、分析与报告的一致决策。以往对大模型代理的评估多聚焦单一环节或单个答案,忽视最终证据包的整体性与内部结构。为此,我们提出RWE-bench,基于MIMIC-IV并源自同行评审的观察性研究,每项任务提供对应研究方案作为参考标准,要求代理在真实数据库中迭代生成树状证据包。我们评估六种LLM(三种开源,三种闭源)在三种代理架构下的表现,采用问题级正确率与端到端任务指标。在162项任务中,任务成功率普遍偏低:最佳代理仅达39.9%,最佳开源模型为30.4%。代理架构影响显著,性能差异超30%。此外,我们实现自动化队列评估方法,快速定位错误并识别代理失效模式。结果表明,当前模型在生成端到端证据包方面仍存在持续局限,高效验证仍是未来重要方向。代码与数据见https://github.com/somewordstoolate/RWE-bench。

原文摘要 · Abstract (English)

Observational studies can yield clinically actionable evidence at scale, but executing them on real-world databases is open-ended and requires coherent decisions across cohort construction, analysis, and reporting. Prior evaluations of LLM agents emphasize isolated steps or single answers, missing the integrity and internal structure of the resulting evidence bundle. To address this gap, we introduce RWE-bench, a benchmark grounded in MIMIC-IV and derived from peer-reviewed observational studies. Each task provides the corresponding study protocol as the reference standard, requiring agents to execute experiments in a real database and iteratively generate tree-structured evidence bundles. We evaluate six LLMs (three open-source, three closed-source) under three agent scaffolds using both question-level correctness and end-to-end task metrics. Across 162 tasks, task success is low: the best agent reaches 39.9%, and the best open-source model reaches 30.4%. Agent scaffolds also matter substantially, causing over 30% variation in performance metrics. Furthermore, we implement an automated cohort evaluation method to rapidly localize errors and identify agent failure modes. Overall, the results highlight persistent limitations in agents' ability to produce end-to-end evidence bundles, and efficient validation remains an important direction for future work. Code and data are available at https://github.com/somewordstoolate/RWE-bench.

大模型评估医学人工智能真实世界证据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。