提出解耦诊断评估框架,分离问诊历史与最终诊断,更公平比较医疗对话模型性能。
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

- 用统一的诊断阅读器评估不同模型生成的问诊历史,固定诊断环节
- 替换诊断器后,诊断准确率变化达2.2-19.0点,排序结果逆转超30%
- 支持可审计的诊断-轨迹-效率三维度评估,适合训练优化和模型对比
评估多轮医疗问诊智能体需判断其通过交互获取病史所提供的诊断支持。现有联合评估让策略同时负责问诊与诊断,导致诊断得分混淆了问诊质量与自身诊断生成能力。本文提出MedDDC-Eval,基于真实病历和在线咨询数据构建的解耦评估基准,对所有策略生成的问诊历史使用相同的冻结共享诊断阅读器,固定终端诊断生成过程,实现跨策略公平比较。该框架报告诊断支持度、信息采集覆盖度与效率。采用大模型辅助语义匹配并进行确定性一对一匹配,确保诊断-轨迹-效率(D/T/E)评分可审计。在固定问诊历史的审计中,将各策略自身的诊断生成器替换为共享阅读器后,诊断F1分数变化2.2-19.0点,记录集与对话集的成对排序有18%和36%发生逆转。为进一步验证下游效用,使用标准组相对策略优化(GRPO)并引入独立训练奖励以优化诊断与轨迹维度。相比Qwen3-32B初始模型,训练后在保留记录与对话测试集上分别提升9.6和4.6分,且任一反馈信号消融均导致两集合得分下降。综上,MedDDC-Eval支持在统一诊断阅读器下进行模型比较,并推动评估驱动的策略开发,同时补充端到端评估以应对诊断生成也属目标能力的情形。
原文摘要 · Abstract (English)
Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through interaction. Yet coupled evaluation lets each policy both elicit the history and generate the terminal diagnosis, so a diagnosis score confounds the elicited history with the policy's own terminal diagnosis generator. We introduce MedDDC-Eval, a diagnosis-decoupled evaluation testbed over held-out cases derived from medical records and online consultations. It applies the same frozen shared diagnostic reader to every policy-elicited history, holding terminal diagnosis generation fixed across policies and enabling comparison under the shared diagnostic reader. It reports diagnostic support, information-acquisition coverage, and efficiency. LLM-assisted semantic matching followed by deterministic one-to-one assignment makes the diagnosis-trajectory-efficiency (D/T/E) scores auditable. In a fixed-history audit across eight policies, replacing each policy's own generator with the shared diagnostic reader shifts diagnosis F1 by 2.2-19.0 points and reverses 18% and 36% of pairwise orderings on the Record and Dialogue splits. To examine downstream utility, we use standard Group Relative Policy Optimization (GRPO) with a separate training-time reward that targets the same diagnosis and trajectory dimensions. Relative to its Qwen3-32B initialization, the trained policy gains 9.6 and 4.6 aggregate-score points on the held-out Record and Dialogue splits, respectively, and ablating either feedback signal reduces the aggregate score on both. Together, MedDDC-Eval supports comparison under a shared diagnostic reader and evaluation-informed policy development, while complementing end-to-end evaluation when terminal diagnosis generation is also part of the target capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。