构建可执行的临床模拟框架,评估医生代理随时间演进的能力。
MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes

- 通过动作触发的模拟门诊事件,动态揭示证据并记录决策过程。
- 700个处理后的病例中,轨迹分析揭示了资源重分配与记忆演进规律。
- 适合研究医疗AI长期学习、跨事件迁移与能力保留的研究者使用。
医生代理正从单轮问答向持续演化的临床决策系统发展。在门诊事件中,它们需获取证据、调用检查与会诊资源,并决定何时形成诊断与管理方案;跨事件间,其行为可能通过记忆、检索、反思等机制更新。现有评估仅部分覆盖此场景:固定输入医疗问答基准仅评分最终答案,多数交互基准仍聚焦单次会诊或固定运行,难以评估事件级决策与跨事件经验的互动。本文提出MedEvoEval,一个基于动作门控的可执行纵向评估框架,将每个原始病例转化为患者、检查与管理者角色视图;证据仅通过合法动作揭示;每例事件记录结构化轨迹,关联观察、动作、最终输出、管理者评分及可选经验写回。我们发布包含700个已处理事件的可运行代码库,含溯源注释、数据模式、事件运行器、评分脚本、配置文件、示例日志、分析代码及轨迹与步级衍生数据。实验表明,事件轨迹暴露了终局评分无法捕捉的过程开销,揭示多学科团队式会诊如何重新分配资源,并支持对记忆成熟、未见迁移、更新阶段响应与反向保留的纵向分析。结果证明,MedEvoEval为评估医生代理是否通过经验提升、转移有用行为并长期保留早期能力提供了实证基础。
原文摘要 · Abstract (English)
Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they acquire evidence, use examination and consultation resources, and decide when to finalize a diagnosis and management plan. Across episodes, their behavior may change through memory, retrieval, reflection, or other update mechanisms. Current evaluations only partially cover this setting. Fixed-input medical QA benchmarks score final answers from complete inputs, whereas many interactive benchmarks still focus on individual encounters or fixed runs, providing limited support for evaluating how episode-level decisions interact with cross-episode experience. We introduce MedEvoEval, an executable longitudinal evaluation framework based on action-gated simulated outpatient episodes. Each source case is converted into role-specific patient, examination, and manager views; evidence is revealed only through valid actions; and each episode records a structured trace that links observations, actions, final outputs, manager scores, and optional experience write-back. We release a runnable E&D artifact with 700 processed episodes, provenance notes, schemas, an episode runner, scoring scripts, configurations, example logs, analysis code, and trajectory- and step-level derivatives. Experiments show that episode traces expose process costs hidden by final-answer scoring, show how MDT-style consultation reallocates resources, and support longitudinal analyses of memory maturation, held-out transfer, update-stage response, and backward retention. Together, these results show that MedEvoEval provides a concrete basis for evaluating whether doctor agents improve through experience, transfer useful behavior, and retain earlier capabilities over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。