arXiv:2512.23440cs.CL2025-12被引 1

用动态对话评估大模型临床推理能力,发现其真实诊疗短板。

ClinDEF: A Dynamic Evaluation Framework for Large Language Models in Clinical Reasoning

  • 基于疾病知识图谱生成动态病历,模拟医生与患者多轮交互。
  • 不仅测诊断准确率,还评估推理效率与诊疗质量细节。
  • 适合医疗AI评测、临床辅助系统开发人员使用。

临床诊断始于医患互动,医生通过反复提问、调整检查和修正鉴别诊断来获取信息。现有LLM评测多聚焦静态问答,无法刻画这一动态过程。虽有研究尝试构建交互式医学框架,但受限于小规模、易污染的数据集,且缺乏细粒度多层级评估。本文提出ClinDEF,一种基于疾病知识图谱的动态评估框架,可生成患者病例并支持大模型医生与自动化患者代理之间的多轮对话。评估协议超越传统诊断准确率,引入细粒度效率分析和基于评分标准的诊断质量评估。实验表明,ClinDEF能有效揭示当前顶尖LLMs在临床推理中的关键缺陷,提供更细致、更具临床意义的评估范式。

原文摘要 · Abstract (English)

Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through patients' response. This dynamic clinical-reasoning process is poorly represented by existing LLM benchmarks that focus on static question-answering. To mitigate these gaps, recent methods explore dynamic medical frameworks involving interactive clinical dialogues. Although effective, they often rely on limited, contamination-prone datasets and lack granular, multi-level evaluation. In this work, we propose ClinDEF, a dynamic framework for assessing clinical reasoning in LLMs through simulated diagnostic dialogues. Grounded in a disease knowledge graph, our method dynamically generates patient cases and facilitates multi-turn interactions between an LLM-based doctor and an automated patient agent. Our evaluation protocol goes beyond diagnostic accuracy by incorporating fine-grained efficiency analysis and rubric-based assessment of diagnostic quality. Experiments show that ClinDEF effectively exposes critical clinical reasoning gaps in state-of-the-art LLMs, offering a more nuanced and clinically meaningful evaluation paradigm.

临床推理大模型评测动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。