构建抗捷径的医学多跳推理基准,揭露大模型真实诊断缺陷
Shattering the Shortcut: A Topology-Regularized Benchmark for Multi-hop Medical Reasoning in LLMs
- 用拓扑规制算法剪除知识图谱中的通用枢纽节点
- 21个大模型在10558题上表现大幅下降,暴露出推理短板
- 适合评估医疗AI深层推理能力,尤其关注临床真实性
尽管大语言模型(LLMs)在标准医学基准上通过单跳事实回忆达到专家水平,但在真实临床场景所需的复杂多跳诊断推理中表现严重不足。主要障碍是‘捷径学习’:模型利用知识图谱中高度连接的通用枢纽节点(如‘炎症’)绕过真实的微病理级联过程。为此,我们提出ShatterMed-QA,一个包含10,558道多跳临床问题的双语基准,用于严格评估深度诊断推理能力。我们的框架采用新颖的$k$-破碎算法构建拓扑规制的医学知识图谱,物理剪除通用枢纽以显式切断逻辑捷径。通过隐式桥接实体掩码和拓扑驱动的困难负样本采样,合成评估案例迫使模型在无表面淘汰依赖的情况下,导航生物合理干扰项。对21个LLMs的全面评估显示,在多跳任务上性能显著下降,尤其是领域专用模型。关键的是,通过检索增强生成(RAG)恢复被掩码证据后,几乎所有模型表现近乎完全恢复,验证了ShatterMed-QA的结构保真度,并证明其有效诊断当前医疗AI的根本推理缺陷。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) achieve expert-level performance on standard medical benchmarks through single-hop factual recall, they severely struggle with the complex, multi-hop diagnostic reasoning required in real-world clinical settings. A primary obstacle is "shortcut learning", where models exploit highly connected, generic hub nodes (e.g., "inflammation") in knowledge graphs to bypass authentic micro-pathological cascades. To address this, we introduce ShatterMed-QA, a bilingual benchmark of 10,558 multi-hop clinical questions designed to rigorously evaluate deep diagnostic reasoning. Our framework constructs a topology-regularized medical Knowledge Graph using a novel $k$-Shattering algorithm, which physically prunes generic hubs to explicitly sever logical shortcuts. We synthesize the evaluation vignettes by applying implicit bridge entity masking and topology-driven hard negative sampling, forcing models to navigate biologically plausible distractors without relying on superficial elimination. Comprehensive evaluations of 21 LLMs reveal massive performance degradation on our multi-hop tasks, particularly among domain-specific models. Crucially, restoring the masked evidence via Retrieval-Augmented Generation (RAG) triggers near-universal performance recovery, validating ShatterMed-QA's structural fidelity and proving its efficacy in diagnosing the fundamental reasoning deficits of current medical AI. Explore the dataset, interactive examples, and full leaderboards at our project website: https://shattermed-qa-web.vercel.app/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。