分离医学模型中的推理与知识能力,发现多数题目只需简单记忆。
Disentangling Reasoning and Knowledge in Medical Large Language Models
- 用PubMedBERT将11个医学问答数据集分为推理和知识两类,准确率达81%
- 仅32.8%的题目需要复杂推理,模型在推理上普遍弱于知识
- 新模型BioMed-R1通过强化学习提升推理鲁棒性,适合临床决策研究者
大型语言模型在医学推理中试图模拟临床诊断思维,但现有基准如MedQA-USMLE、MedMCQA和PubMedQA常混淆推理与事实记忆。本文通过PubMedBERT分类器将11个生物医学问答基准拆分为推理与知识子集,准确率达81%,接近人类水平。分析显示,仅有32.8%的问题需要复杂推理。评估了多个生物医学模型(HuatuoGPT-o1、MedReason、m1)和通用模型(DeepSeek-R1、o4-mini、Qwen3),均发现知识与推理性能存在明显差距:例如,HuatuoGPT-o1知识得分56.9,推理仅44.8。在误导性推理的对抗测试中,生物医学模型表现急剧下降,而更大或经强化学习训练的通用模型更具鲁棒性。为此,我们基于推理密集样本,对BioMed-R1进行微调与强化学习训练,其在同规模模型中表现最优。未来可通过引入临床病例报告及对抗与回溯场景进一步提升性能。
原文摘要 · Abstract (English)
Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reasoning with factual recall. We address this by separating 11 biomedical QA benchmarks into reasoning- and knowledge-focused subsets using a PubMedBERT classifier that reaches 81 percent accuracy, comparable to human performance. Our analysis shows that only 32.8 percent of questions require complex reasoning. We evaluate biomedical models (HuatuoGPT-o1, MedReason, m1) and general-domain models (DeepSeek-R1, o4-mini, Qwen3), finding consistent gaps between knowledge and reasoning performance. For example, HuatuoGPT-o1 scores 56.9 on knowledge but only 44.8 on reasoning. In adversarial tests where models are misled with incorrect initial reasoning, biomedical models degrade sharply, while larger or RL-trained general models show more robustness. To address this, we train BioMed-R1 using fine-tuning and reinforcement learning on reasoning-heavy examples. It achieves the strongest performance among similarly sized models. Further gains may come from incorporating clinical case reports and training with adversarial and backtracking scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。