arXiv:2510.20001cs.CLcs.AI2025-10被引 1

提出双维度框架,让大模型临床决策评估更贴近真实医疗场景。

Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs

  • 从临床背景和问题类型双维度重构医疗决策任务
  • 现有数据集评估标准单一,真实场景下模型表现显著下降
  • 强调效率与可解释性,适合临床落地研究者参考

大型语言模型(LLMs)在临床应用中展现潜力,但通常依赖如MedQA等简化问答数据集进行评估。这类数据集未能充分反映真实临床决策的复杂性。为此,我们提出一个统一范式,从临床背景和临床问题两个维度刻画决策任务。随着背景与问题越接近真实医疗环境,任务难度越高。我们系统梳理了现有数据集与基准测试在这两个维度上的分布,并回顾了应对临床决策的方法,包括训练期与推理期技术,总结其适用场景。此外,将评估维度从准确率扩展至效率与可解释性。最后,指出了当前关键挑战。该范式有助于澄清假设、标准化比较,并指导更具临床意义的LLM研发。

原文摘要 · Abstract (English)

Large language models (LLMs) show promise for clinical use. They are often evaluated using datasets such as MedQA. However, Many medical datasets, such as MedQA, rely on simplified Question-Answering (Q\A) that underrepresents real-world clinical decision-making. Based on this, we propose a unifying paradigm that characterizes clinical decision-making tasks along two dimensions: Clinical Backgrounds and Clinical Questions. As the background and questions approach the real clinical environment, the difficulty increases. We summarize the settings of existing datasets and benchmarks along two dimensions. Then we review methods to address clinical decision-making, including training-time and test-time techniques, and summarize when they help. Next, we extend evaluation beyond accuracy to include efficiency, explainability. Finally, we highlight open challenges. Our paradigm clarifies assumptions, standardizes comparisons, and guides the development of clinically meaningful LLMs.

临床决策大模型评估医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。