arXiv:2603.25821cs.CLcs.AI2026-03

用模拟医患对话评估医疗AI,比传统考试更贴近真实诊疗流程。

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

  • 构建多步医患对话场景,测试AI采集病史、诊断与治疗建议能力。
  • 基于254例合成病例评估,覆盖诊断、鉴别诊断、安全处理等6个维度。
  • 适合研究临床推理过程,也用于模型迭代中的行为监控与对比。

我们提出Doctorina MedBench,一个基于模拟医患对话的代理式医疗AI评估框架。不同于依赖标准化试题的传统基准,该框架模拟多步骤临床对话,要求AI系统收集病史、分析合成附件、形成鉴别诊断并提供诊疗建议。性能在诊断、鉴别诊断、治疗、安全事件处理及对话行为等任务层面分别评估,并采用D.O.T.S.框架(诊断、观察/检查、治疗、步骤数)进行综合汇总。框架支持开发过程中的行为变化监测、安全导向案例测试、类别化合成场景采样及版本间回归对比。本研究分析了254例由医生撰写的合成临床案例(来自261个尝试生成的标识符),评估指标旨在促进交互式医疗AI系统的比较研究与临床推理流程分析。结果表明,模拟临床对话可作为传统考试类基准的补充评估方式,但未验证其独立临床有效性、实际疗效或现实部署可行性。

原文摘要 · Abstract (English)

We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions. Unlike traditional medical benchmarks that rely on solving standardized test questions, the proposed approach models a multi-step clinical dialogue in which an AI system must collect medical history, analyze available synthetic attachments when present, formulate differential diagnoses, and provide diagnostic and management recommendations. System performance is evaluated across separate task-level domains, including diagnosis, differential diagnosis, treatment, safety-critical condition handling, and dialogue-step behavior; the broader D.O.T.S. framework is used as a supplementary summary for diagnosis, observations/investigations, treatment, and step count. The framework also supports testing and quality-monitoring workflows intended to identify changes in model behavior during development. It supports safety-oriented cases, category-based sampling of synthetic clinical scenarios, and regression-style comparisons across system versions. In the reported study, the analyzed paired complete-case cohort consisted of 254 physician-authored synthetic clinical cases retained from 261 attempted case identifiers. The evaluation metrics are intended for comparative research on interactive medical AI systems and for studying clinical reasoning workflows in synthetic dialogue settings. Our results suggest that simulated clinical dialogue can provide a complementary assessment setting to traditional examination-style benchmarks, while the reported findings do not establish independent clinical validity, clinical effectiveness, or readiness for real-world deployment.

医疗AI对话评估临床推理合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。