用智能检索增强推理,让不同医生模型更一致地给出可靠诊断答案
Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering
- 让大模型分步检索医学知识并生成结构化报告,统一决策依据
- 模型间答案差异缩小一半(熵从0.48降至0.13),正确率稳定提升至81%
- 适合关注临床部署中模型可靠性与一致性研究的医疗AI从业者
智能检索增强推理系统在临床决策支持中日益普及,通过迭代检索领域知识并生成结构化报告来指导大语言模型(LLMs)作答。我们评估了34个LLMs在169个专家标注的放射科问题上的表现,对比零样本推理与放射科专用多步智能检索条件下的结果。在后一种条件下,所有模型均接收相同来源的结构化证据报告。结果显示,智能推理显著降低模型间决策分散度(中位熵从0.48降至0.13),提升跨模型正确性稳定性(平均正确率从0.74升至0.81),多数意见一致性显著增强(P<0.001)。尽管高一致不保证正确,但共识强度与正确性仍高度相关(零样本时ρ=0.88,智能推理时ρ=0.87)。响应冗长性与正确性无关。在572个错误输出中,72%被评估为中到高临床严重程度,但评分者间一致性较低(κ=0.02)。这表明智能检索可提升决策集中度、共识力和跨模型鲁棒性,提示仅靠准确率或一致率不足以衡量可靠性,需结合稳定性、鲁棒性及临床影响综合评估。
原文摘要 · Abstract (English)
Agentic retrieval-augmented reasoning pipelines are increasingly used to structure how large language models (LLMs) incorporate external evidence in clinical decision support. These systems iteratively retrieve curated domain knowledge and synthesize it into structured reports before answer selection. Although such pipelines can improve performance, their impact on reliability under model variability remains unclear. In real-world deployment, heterogeneous models may align, diverge, or synchronize errors in ways not captured by accuracy. We evaluated 34 LLMs on 169 expert-curated publicly available radiology questions, comparing zero-shot inference with a radiology-specific multi-step agentic retrieval condition in which all models received identical structured evidence reports derived from curated radiology knowledge. Agentic inference reduced inter-model decision dispersion (median entropy 0.48 vs. 0.13) and increased robustness of correctness across models (mean 0.74 vs. 0.81). Majority consensus also increased overall (P<0.001). Consensus strength and robust correctness remained correlated under both strategies (\r{ho}=0.88 for zero-shot; \r{ho}=0.87 for agentic), although high agreement did not guarantee correctness. Response verbosity showed no meaningful association with correctness. Among 572 incorrect outputs, 72% were associated with moderate or high clinically assessed severity, although inter-rater agreement was low (\k{appa}=0.02). Agentic retrieval therefore was associated with more concentrated decision distributions, stronger consensus, and higher cross-model robustness of correctness. These findings suggest that evaluating agentic systems through accuracy or agreement alone may not always be sufficient, and that complementary analyses of stability, cross-model robustness, and potential clinical impact are needed to characterize reliability under model variability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。