构建医疗对话评估新基准,专测大模型问诊能力。
MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language Models
- 用多智能体生成5200个虚拟病例,模拟真实问诊流程。
- 专家审核6万余条评分标准,确保诊断逻辑科学可靠。
- 发现现有模型问诊能力严重不足,需改进对话架构。
医疗对话AI在构建更安全、高效的医疗对话系统中至关重要。然而,现有评估框架对医疗大模型的信息收集与诊断推理能力缺乏严格检验。为此,我们提出MedDialogRubrics,一个包含5,200个合成患者病例和超过60,000条由大模型生成、经临床专家修正的细粒度评价标准的新基准,专门用于评估大模型的多轮诊断能力。该框架采用多智能体系统,基于疾病知识合成真实患者病历与主诉,不依赖真实电子病历,规避隐私与数据治理风险。我们设计了受限于基础医学事实并具备动态纠错机制的患者智能体,持续检测并修正幻觉,保障对话内部一致性和临床合理性。此外,提出基于结构化大模型与专家标注的评分标准生成流程,检索循证医学(EBM)指南,利用拒绝采样提取每例病例的“必须提问”项。对主流模型的全面评估显示,当前模型在多个维度仍面临巨大挑战。结果表明,提升医疗对话性能需在对话管理架构上突破,而不仅是对基础模型的微调。
原文摘要 · Abstract (English)
Medical conversational AI (AI) plays a pivotal role in the development of safer and more effective medical dialogue systems. However, existing benchmarks and evaluation frameworks for assessing the information-gathering and diagnostic reasoning abilities of medical large language models (LLMs) have not been rigorously evaluated. To address these gaps, we present MedDialogRubrics, a novel benchmark comprising 5,200 synthetically constructed patient cases and over 60,000 fine-grained evaluation rubrics generated by LLMs and subsequently refined by clinical experts, specifically designed to assess the multi-turn diagnostic capabilities of LLM. Our framework employs a multi-agent system to synthesize realistic patient records and chief complaints from underlying disease knowledge without accessing real-world electronic health records, thereby mitigating privacy and data-governance concerns. We design a robust Patient Agent that is limited to a set of atomic medical facts and augmented with a dynamic guidance mechanism that continuously detects and corrects hallucinations throughout the dialogue, ensuring internal coherence and clinical plausibility of the simulated cases. Furthermore, we propose a structured LLM-based and expert-annotated rubric-generation pipeline that retrieves Evidence-Based Medicine (EBM) guidelines and utilizes the reject sampling to derive a prioritized set of rubric items ("must-ask" items) for each case. We perform a comprehensive evaluation of state-of-the-art models and demonstrate that, across multiple assessment dimensions, current models face substantial challenges. Our results indicate that improving medical dialogue will require advances in dialogue management architectures, not just incremental tuning of the base-model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。