DeepSeek R1在医学诊断中表现接近专家,但存在偏见和推理过长等缺陷。
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
- 用100个真实病例测试模型,评估其临床推理能力
- 诊断准确率达93%,但7例错误暴露系统性缺陷
- 响应越短越可靠,适合做决策辅助而非独立诊断
将大型语言模型(如DeepSeek R1)应用于医疗领域需严格评估其推理是否符合临床规范。本研究基于100个MedQA临床案例,评估DeepSeek R1的医学推理与专家模式的一致性。模型诊断准确率达到93%,展现出系统性临床判断能力,包括鉴别诊断、基于指南的治疗选择及个体化因素整合。然而,对7个错误案例的分析揭示了持续存在的局限:锚定偏差、冲突数据协调困难、备选方案探索不足、过度思考、知识空白以及过早优先确定最终治疗方案。关键发现是推理长度与准确性正相关——响应字符数少于5000时更可靠,表明冗长解释可能反映不确定或错误合理化。尽管DeepSeek R1具备基础临床推理能力,但重复出现的缺陷提示需在偏见缓解、知识更新和结构化推理框架方面改进。研究强调,尽管大模型有潜力通过人工推理辅助医疗决策,但仍需领域特定验证、可解释性保障及置信度指标(如响应长度阈值),以确保实际应用中的可靠性。
原文摘要 · Abstract (English)
Integrating large language models (LLMs) like DeepSeek R1 into healthcare requires rigorous evaluation of their reasoning alignment with clinical expertise. This study assesses DeepSeek R1's medical reasoning against expert patterns using 100 MedQA clinical cases. The model achieved 93% diagnostic accuracy, demonstrating systematic clinical judgment through differential diagnosis, guideline-based treatment selection, and integration of patient-specific factors. However, error analysis of seven incorrect cases revealed persistent limitations: anchoring bias, challenges reconciling conflicting data, insufficient exploration of alternatives, overthinking, knowledge gaps, and premature prioritization of definitive treatment over intermediate care. Crucially, reasoning length correlated with accuracy - shorter responses (<5,000 characters) were more reliable, suggesting extended explanations may signal uncertainty or rationalization of errors. While DeepSeek R1 exhibits foundational clinical reasoning capabilities, recurring flaws highlight critical areas for refinement, including bias mitigation, knowledge updates, and structured reasoning frameworks. These findings underscore LLMs' potential to augment medical decision-making through artificial reasoning but emphasize the need for domain-specific validation, interpretability safeguards, and confidence metrics (e.g., response length thresholds) to ensure reliability in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。