评估大模型在临床问答中的表现,发现检索增强能显著提升准确率。
Evaluating Large Language Models for Evidence-Based Clinical Question Answering
- 构建多源临床指南与系统综述数据集,测试模型对结构化和叙述性问题的回答能力。
- 模型在结构化指南上准确率达90%,在综述类问题上为60%~70%,且引用越多越易答对。
- 引入相关文献摘要可使错误答案准确率提升至79%,适合医疗AI可靠性研究者参考。
大语言模型在生物医学和临床应用中取得显著进展,促使对其回答复杂、基于证据的临床问题能力进行严格评估。我们构建了一个多源基准数据集,涵盖Cochrane系统综述和临床指南,包括美国心脏协会的结构化建议及保险公司使用的叙述性指导。使用GPT-4o-mini和GPT-5测试发现,模型在结构化指南推荐上的准确率最高(90%),而在叙述性指南和系统综述问题上较低(60%–70%)。准确率与底层系统综述的引用次数强相关:每引用量翻倍,正确回答的几率约提升30%。模型在提供上下文时具备中等证据质量推理能力。引入检索增强提示后,若提供黄金来源摘要,原错误项准确率升至0.79;提供前3篇按语义相关性排序的PubMed摘要,准确率提升至0.23;随机摘要则导致准确率下降至0.10(温度变化范围内)。该现象在GPT-4o-mini中同样成立,表明源文本清晰度和精准检索比模型规模更关键。总体而言,结果揭示了大模型在循证临床问答中的潜力与局限,检索增强是提升事实准确性的重要策略,而按专科与题型分层评估仍为理解知识获取现状和定位模型性能的关键。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated substantial progress in biomedical and clinical applications, motivating rigorous evaluation of their ability to answer nuanced, evidence-based questions. We curate a multi-source benchmark drawing from Cochrane systematic reviews and clinical guidelines, including structured recommendations from the American Heart Association and narrative guidance used by insurers. Using GPT-4o-mini and GPT-5, we observe consistent performance patterns across sources and clinical domains: accuracy is highest on structured guideline recommendations (90%) and lower on narrative guideline and systematic review questions (60--70%). We also find a strong correlation between accuracy and the citation count of the underlying systematic reviews, where each doubling of citations is associated with roughly a 30% increase in the odds of a correct answer. Models show moderate ability to reason about evidence quality when contextual information is supplied. When we incorporate retrieval-augmented prompting, providing the gold-source abstract raises accuracy on previously incorrect items to 0.79; providing top 3 PubMed abstracts (ranked by semantic relevance) improves accuracy to 0.23, while random abstracts reduce accuracy (0.10, within temperature variation). These effects are mirrored in GPT-4o-mini, underscoring that source clarity and targeted retrieval -- not just model size -- drive performance. Overall, our results highlight both the promise and current limitations of LLMs for evidence-based clinical question answering. Retrieval-augmented prompting emerges as a useful strategy to improve factual accuracy and alignment with source evidence, while stratified evaluation by specialty and question type remains essential to understand current knowledge access and to contextualize model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。