让AI在医学问答中边查边反思,答案更准更有依据。
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
- 分三阶段:先优化检索词,再批量查文献,最后带引用生成答案
- 在PubMedQA上达78.32%准确率,略超人类专家
- 适合临床医生和科研人员用,还能控制计算成本
可信的生物医学问答系统不仅需提供准确答案,还须以当前可验证的证据加以支持。现有检索增强方法缺乏对劣质查询的迭代优化机制,而自省方法仅在完整检索后才启动。为此,我们提出PubMed Reasoner,一种由三个阶段组成的生物医学QA智能体:自我批判式查询优化基于部分(元数据)检索评估医学主题词的覆盖度、匹配度与冗余性,以改进PubMed查询;反思式检索以批次方式处理文献,直至收集到足够证据;证据锚定式回答生成则输出附带明确引用的答案。采用GPT-4o作为主干模型的PubMed Reasoner在PubMedQA上达到78.32%准确率,略高于人类专家,并在MMLU临床知识测试中持续提升表现。此外,大模型作为评判者的结果显示,我们的回答在推理合理性、证据支撑性、临床相关性和可信度方面均更受青睐。通过在权威来源上实现检索优先的动态推理,该方法为临床医生和生物医学研究人员提供了实用辅助,同时有效控制计算与令牌开销。
原文摘要 · Abstract (English)
Trustworthy biomedical question answering (QA) systems must not only provide accurate answers but also justify them with current, verifiable evidence. Retrieval-augmented approaches partially address this gap but lack mechanisms to iteratively refine poor queries, whereas self-reflection methods kick in only after full retrieval is completed. In this context, we introduce PubMed Reasoner, a biomedical QA agent composed of three stages: self-critic query refinement evaluates MeSH terms for coverage, alignment, and redundancy to enhance PubMed queries based on partial (metadata) retrieval; reflective retrieval processes articles in batches until sufficient evidence is gathered; and evidence-grounded response generation produces answers with explicit citations. PubMed Reasoner with a GPT-4o backbone achieves 78.32% accuracy on PubMedQA, slightly surpassing human experts, and showing consistent gains on MMLU Clinical Knowledge. Moreover, LLM-as-judge evaluations prefer our responses across: reasoning soundness, evidence grounding, clinical relevance, and trustworthiness. By orchestrating retrieval-first reasoning over authoritative sources, our approach provides practical assistance to clinicians and biomedical researchers while controlling compute and token costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。