arXiv:2605.21807cs.CL2026-05

测试大模型在罕见病例中超越指南的推理能力,发现需检索外部证据才能提升准确率。

When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

论文配图:When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering
图 1 · 摘自论文原文
  • 构建自由文本问答的检索基准,评估模型处理非指南覆盖病例的能力。
  • 最先进模型仅56%正确率,引入文献检索后提升至82%。
  • 适合医疗AI研发者和需要真实世界推理能力的研究者使用。

在多个医学领域,临床实践依赖循证指南来规范诊断与治疗路径。然而,这些路径难以覆盖实际诊疗中的长尾情况。当前大多数医学大模型主要通过参数记忆常见、指南相关的知识,评估也多集中于对这类内容的回忆与推理,常采用选择题形式。由于医学实践中证据推理至关重要,依赖记忆不可靠且不现实。为此,我们提出OGCaReBench——一个以自由文本回答为核心的检索型评估基准,专门用于评测大模型在需突破常规指南的临床问题上的表现。该基准数据来自已发表的病例报告,并经医学专家验证,包含需要长篇自由作答的临床问题,系统性地支持对罕见病例下开放式医学推理的评估。实验表明,即使最佳基线模型(GPT-5.2)也只能正确回答56%的问题,专用模型表现更低,仅为42%。通过引入检索医学文献,模型性能最高可提升至82%(使用GPT-5.2),凸显了证据接地在真实医疗推理任务中的关键作用。本研究为通用与医学大模型在复杂临床场景下的可靠应答能力评估与改进奠定了基础。

原文摘要 · Abstract (English)

Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall short for the long tail of real-world care not covered by guidelines. Most medical large language models (LLMs), however, are trained to encode common, guideline-focused medical knowledge in their parameters. Current evaluations test models primarily on recalling and reasoning with this memorized content, often in multiple-choice settings. Given the fundamental importance of evidence-based reasoning in medicine, it is neither feasible nor reliable to depend on memorization in practice. To address this gap, we introduce OGCaReBench, a free-form retrieval-focused benchmark aimed at evaluating LLMs at answering clinical questions that require going beyond typical guidelines. Extracted from published medical case reports and validated by medical experts, OGCaReBench contains long-form clinical questions requiring free-text answers, providing a systematic framework for assessing open-ended medical reasoning in rare, case-based scenarios. Our experiments reveal that even the best-performing baseline (GPT-5.2) correctly answers only 56% of our benchmark with specialized models only reaching 42%. Augmenting models with retrieved medical articles improves this performance to up to 82% (using GPT-5.2) highlighting the importance of evidence-grounding for real-world medical reasoning tasks. This work thus establishes a foundation for benchmarking and advancing both general-purpose and medical LLMs to produce reliable answers in challenging clinical contexts.

医学AI检索增强罕见病大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。