arXiv:2507.02975cs.LGcs.IR2025-07

评估大模型生物医学回答是否基于真实证据的框架

Introducing Answered with Evidence -- a framework for evaluating whether LLM responses to biomedical questions are founded in evidence

  • 构建多源检索增强生成系统,融合文献与新型证据库
  • 仅44%问题获文献支持,新证据源提升至50%,合并超70%可信赖
  • 适合医疗AI验证、临床决策辅助系统开发者参考

大语言模型在生物医学问答中的应用日益广泛,但其回答的准确性与证据支持备受关注。为此,我们提出Answered with Evidence框架,评估LLM回答是否基于科学文献。通过分析数千个医生提交的问题,采用对比管道:(1) Alexandria(原Atropos证据库),基于新型观察性研究的检索增强生成系统;(2) 两个基于PubMed的RAG系统(System和Perplexity)。结果表明,基于PubMed的系统对约44%的问题提供有证据支持的回答,而新型证据源则覆盖约50%。两者结合使超过70%的生物医学问题可获得可靠答案。随着LLM总结科学内容能力提升,要最大化其价值,需具备准确检索已发表及定制生成证据的能力,或实时生成证据。

原文摘要 · Abstract (English)

The growing use of large language models (LLMs) for biomedical question answering raises concerns about the accuracy and evidentiary support of their responses. To address this, we present Answered with Evidence, a framework for evaluating whether LLM-generated answers are grounded in scientific literature. We analyzed thousands of physician-submitted questions using a comparative pipeline that included: (1) Alexandria, fka the Atropos Evidence Library, a retrieval-augmented generation (RAG) system based on novel observational studies, and (2) two PubMed-based retrieval-augmented systems (System and Perplexity). We found that PubMed-based systems provided evidence-supported answers for approximately 44% of questions, while the novel evidence source did so for about 50%. Combined, these sources enabled reliable answers to over 70% of biomedical queries. As LLMs become increasingly capable of summarizing scientific content, maximizing their value will require systems that can accurately retrieve both published and custom-generated evidence or generate such evidence in real time.

生物医学证据评估RAGLLM验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。