让医疗对话代理自动学习并改进诊疗流程,提升诊断准确率与安全性。
MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

- 通过四类知识库动态更新临床、流程、符号与视觉信息,实现无微调自进化。
- 在MIMIC-IV数据上诊断准确率提升7.81%,治疗意图覆盖率达70.67%。
- 适合医疗AI研究者及需要高可靠诊疗系统的开发者使用。
交互式临床代理在部分可观测条件下运行,可靠诊疗依赖于基于证据的安全互动。现有代理难以将经验转化为具有明确来源和权威性的可复用流程知识。为此,我们提出MediSkill-Evo,一种无需微调主干模型的流程约束自演化方法。它通过四类知识库(临床、流程、符号、视觉)在类型特定验证与作用范围规则下更新知识,并利用流程约束偏好机制,将验证后的知识转化为行动,以证据为依据并优先安全决策。我们在300个基于MIMIC-IV的FullChain病例、180种硬隔离条件(涵盖六项流程义务)以及100个包含NEJM图像的多模态病例上评估。在Qwen FullChain上,MediSkill-Evo相比最优前代代理,诊断准确率提升7.81%,治疗意图覆盖提高70.67%,关键失误减少43.04%。在压力测试中,复合流程表现提升7.77%,必要动作完成率提高12.41%,且在患者事实、时间证据与分诊警报恢复方面表现更强,未出现控制器评分错误。在多模态NEJM诊断任务中,结合可选MedSAM定位,诊断准确率提升2.56%,核心得分提高18.96%。代码已公开。
原文摘要 · Abstract (English)
Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existing agents struggle to convert experience into reusable process knowledge with explicit provenance and authority. To address this gap, we introduce MediSkill-Evo, which self-evolves governed process knowledge without fine-tuning the backbone. It realizes this self-evolution by updating clinical, process, symbolic, and visual knowledge in four typed banks under type-specific validation and scope rules. The Process-Constrained Preference Harness then turns validated knowledge into action by grounding candidates in evidence and prioritizing safer decisions. We evaluate on 300 MIMIC-IV-derived FullChain encounters, 180 hard-isolation conditions covering six process obligations, and 100 multimodal NEJM image-diagnosis cases. On Qwen FullChain, MediSkill-Evo improves diagnosis accuracy by 7.81% and treatment-intent coverage by 70.67% over the best-performing prior agent, while reducing critical failures by 43.04%. Under stress, it improves the stress-process composite by 7.77% and required-action completion by 12.41% over the best-performing agent for each metric, with stronger patient-fact, temporal-evidence, and triage-red-flag recovery and no controller-scored errors in unavailable-evidence, treatment, and triage safety checks. On multimodal NEJM diagnosis, MediSkill-Evo with optional MedSAM localization improves diagnosis accuracy by 2.56% and core score by 18.96% over the best-performing memory agent. Code is available at https://anonymous.4open.science/r/mediskill-evo_anonymous-68E7.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。