arXiv:2509.21933cs.CLcs.AI2025-09被引 9

86.3%的临床大模型在思维链提示下性能下降,揭示其可解释性与可靠性矛盾。

Why Chain of Thought Fails in Clinical Text Understanding

  • 首次大规模测试思维链在临床文本中的表现,覆盖95个模型、87项任务。
  • 86.3%模型在思维链设置下性能下降,越弱模型受损越严重。
  • 发现思维链虽提升解释性,却可能降低临床判断可靠性,需谨慎使用。

大型语言模型(LLMs)正越来越多地应用于临床医疗,该领域对准确性和透明推理至关重要。思维链(CoT)提示通过引导分步推理,在多个任务中提升了性能和可解释性。然而,其在电子健康记录(EHRs)等临床场景中的有效性仍不明确,而EHRs通常冗长、碎片化且含噪声。本文首次开展大规模系统性研究,评估95个先进LLMs在87项真实临床任务上的表现,涵盖9种语言和8类任务类型。结果表明,与其它领域相反,86.3%的模型在思维链设置下出现持续性能下降。更强大的模型相对稳健,而较弱模型则遭受显著退化。通过细粒度分析推理长度、医学概念对齐及错误模式,并结合大模型判别与临床专家评估,我们揭示了思维链在临床语境中失效的系统性规律。这一发现凸显了一个关键悖论:思维链增强了可解释性,却可能削弱临床任务的可靠性。本研究为临床大模型的推理策略提供了实证基础,强调需发展更透明可信的方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being applied to clinical care, a domain where both accuracy and transparent reasoning are critical for safe and trustworthy deployment. Chain-of-thought (CoT) prompting, which elicits step-by-step reasoning, has demonstrated improvements in performance and interpretability across a wide range of tasks. However, its effectiveness in clinical contexts remains largely unexplored, particularly in the context of electronic health records (EHRs), the primary source of clinical documentation, which are often lengthy, fragmented, and noisy. In this work, we present the first large-scale systematic study of CoT for clinical text understanding. We assess 95 advanced LLMs on 87 real-world clinical text tasks, covering 9 languages and 8 task types. Contrary to prior findings in other domains, we observe that 86.3\% of models suffer consistent performance degradation in the CoT setting. More capable models remain relatively robust, while weaker ones suffer substantial declines. To better characterize these effects, we perform fine-grained analyses of reasoning length, medical concept alignment, and error profiles, leveraging both LLM-as-a-judge evaluation and clinical expert evaluation. Our results uncover systematic patterns in when and why CoT fails in clinical contexts, which highlight a critical paradox: CoT enhances interpretability but may undermine reliability in clinical text tasks. This work provides an empirical basis for clinical reasoning strategies of LLMs, highlighting the need for transparent and trustworthy approaches.

大模型临床推理思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。