arXiv:2508.00743cs.CLcs.AI2025-08被引 19

多步检索推理提升医学影像问答准确率

Multi-step retrieval and reasoning improves radiology question answering with large language models

  • 设计多步检索与推理框架,分步获取关键影像信息
  • 小模型提升超20%,减少9.4%幻觉,46%案例获临床相关上下文
  • 适合医疗AI研究者及临床辅助系统开发者使用

放射科临床决策日益依赖人工智能,尤其是大语言模型(LLMs)。然而,传统放射科问答(QA)的检索增强生成(RAG)系统通常依赖单步检索,难以应对复杂临床推理任务。本文提出放射科检索与推理(RaR)框架,通过多步检索与推理提升诊断准确性、事实一致性和临床可靠性。评估涵盖25种不同架构、参数规模(0.5B至>670B)和训练范式的模型,使用104个专家标注的放射科问题(来自RSNA-RadioQA和ExtendedQA数据集),并额外在65个真实放射科考试题组成的内部未见数据集上测试。结果表明,相比零样本提示和传统在线RAG,RaR显著提升平均诊断准确率,尤其在小模型中增益最大;超过2000亿参数的大模型改进不足2%。此外,RaR将幻觉率降至均值9.4%,在46%案例中成功检索到临床相关上下文,显著增强事实依据。即使经过临床微调的模型(如MedGemma-27B)也从中受益,说明检索对已有领域知识仍有价值。所有数据集、代码和完整框架均已公开,支持开放研究与临床转化。

原文摘要 · Abstract (English)

Clinical decision-making in radiology increasingly benefits from artificial intelligence (AI), particularly through large language models (LLMs). However, traditional retrieval-augmented generation (RAG) systems for radiology question answering (QA) typically rely on single-step retrieval, limiting their ability to handle complex clinical reasoning tasks. Here we propose radiology Retrieval and Reasoning (RaR), a multi-step retrieval and reasoning framework designed to improve diagnostic accuracy, factual consistency, and clinical reliability of LLMs in radiology question answering. We evaluated 25 LLMs spanning diverse architectures, parameter scales (0.5B to >670B), and training paradigms (general-purpose, reasoning-optimized, clinically fine-tuned), using 104 expert-curated radiology questions from previously established RSNA-RadioQA and ExtendedQA datasets. To assess generalizability, we additionally tested on an unseen internal dataset of 65 real-world radiology board examination questions. RaR significantly improved mean diagnostic accuracy over zero-shot prompting and conventional online RAG. The greatest gains occurred in small-scale models, while very large models (>200B parameters) demonstrated minimal changes (<2% improvement). Additionally, RaR retrieval reduced hallucinations (mean 9.4%) and retrieved clinically relevant context in 46% of cases, substantially aiding factual grounding. Even clinically fine-tuned models showed gains from RaR (e.g., MedGemma-27B), indicating that retrieval remains beneficial despite embedded domain knowledge. These results highlight the potential of RaR to enhance factuality and diagnostic accuracy in radiology QA, warranting future studies to validate their clinical utility. All datasets, code, and the full RaR framework are publicly available to support open research and clinical translation.

医学影像大模型检索增强临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。