arXiv:2608.22713cs.CL2026-08中稿 · IEEE HealthCom 202…

从病历报告构建可评估的多模态问诊对话,检验大模型真实诊断能力

A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports

论文配图:A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports
图 1 · 摘自论文原文
  • 基于病历报告生成逐步推进的多模态问诊对话,确保每步回答有据可依
  • 实测诊断准确率F1达0.99,推理质量得分4.79,远超主流大模型表现
  • 适合评估医疗大模型的推理与影像理解能力,推动临床智能系统发展

临床诊断需逐步整合患者病史、体格检查、实验室结果、医学影像及诊断性检查。然而,现有多数多模态医疗评测仅评估固定输入或最终答案,而全交互式诊断代理常混淆证据选择与解释。本文提出一种源基框架,可从内科病例报告中构建渐进式多模态问诊对话,并设计评估策略,用于衡量多模态大语言模型(MLLMs)在最终诊断、诊断推理及影像发现解读方面的能力。在24例内科病例上的评估显示,该框架能准确生成参考对话,诊断F1为0.99,推理质量得分为5分制中的4.79。对o4-mini和Claude Haiku 4.5两个前沿MLLM的评估结果显示,其推理质量得分分别为2.75和2.50,诊断、推理与影像发现的F1分数显著偏低。结果表明,流畅的回答不等于基于证据的临床推理,凸显了所提框架在评估多模态诊断推理中的价值。

原文摘要 · Abstract (English)

Clinical diagnosis requires progressive integration of patient history, physical examination, laboratory findings, medical images, and diagnostic-informative tests. However, most multimodal medical benchmarks evaluate fixed inputs or endpoint answers, while fully interactive diagnostic agents conflate evidence selection with evidence interpretation. We present a source-grounded framework to construct progressive multimodal diagnostic dialogues from case reports and an evaluation strategy for assessing MLLMs on final diagnosis, diagnostic reasoning, and image-finding interpretation. Evaluation on 24 internal medicine case reports showed that our framework can accurately convert case reports into reference dialogues, achieving a diagnosis F1 of 0.99 and a reasoning-quality score of 4.79 out of 5. Evaluation on two frontier MLLMs (o4-mini and Claude Haiku 4.5) achieved reasoning-quality scores of 2.75 and 2.50, respectively, with substantially lower diagnosis, reasoning, and image-finding F1 scores. The results demonstrate that fluent responses do not necessarily reflect evidence-grounded clinical reasoning and highlight the utility of the proposed framework for evaluating multimodal diagnostic reasoning.

多模态诊断临床推理大模型评估病历生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。