arXiv:2412.15748cs.CLcs.AI2024-12被引 15

剖析医学大模型的推理机制,提升临床可信度。

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models

  • 提出医疗大模型推理行为的分类框架
  • 揭示模型内部低层推理逻辑,增强可解释性
  • 为医生与工程师提供诊断辅助工具

尽管大型语言模型(LLMs)在医疗领域广泛应用,但对其推理行为的研究仍严重不足。本文强调理解推理行为的重要性,因其等同于医疗场景中的可解释AI(XAI)。我们拓展并重构了推理行为的概念,将其应用于医疗大模型上下文。系统综述并分类当前最先进的医疗大模型推理建模与评估方法。同时,提出理论框架,帮助医疗从业者或机器学习工程师洞察模型内部的低层推理操作。最后,梳理了构建大型推理模型面临的关键开放挑战。提升医疗机器学习模型的透明度与可信度,将促进其在临床实践中的采纳、应用及进一步发展。

原文摘要 · Abstract (English)

Background: Despite the current ubiquity of Large Language Models (LLMs) across the medical domain, there is a surprising lack of studies which address their reasoning behaviour. We emphasise the importance of understanding reasoning behaviour as opposed to high-level prediction accuracies, since it is equivalent to explainable AI (XAI) in this context. In particular, achieving XAI in medical LLMs used in the clinical domain will have a significant impact across the healthcare sector. Results: Therefore, in this work, we adapt the existing concept of reasoning behaviour and articulate its interpretation within the specific context of medical LLMs. We survey and categorise current state-of-the-art approaches for modeling and evaluating reasoning reasoning in medical LLMs. Additionally, we propose theoretical frameworks which can empower medical professionals or machine learning engineers to gain insight into the low-level reasoning operations of these previously obscure models. We also outline key open challenges facing the development of Large Reasoning Models. Conclusion: The subsequent increased transparency and trust in medical machine learning models by clinicians as well as patients will accelerate the integration, application as well as further development of medical AI for the healthcare system as a whole.

医疗AI可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。