arXiv:2603.10677cs.AIcs.CL2026-03被引 1

让AI像医生一样持续学习诊断,自动优化判断能力。

Emulating Clinician Cognition via Self-Evolving Deep Clinical Research

  • 构建可自我进化诊断系统,模拟医生动态问诊与经验积累
  • 在MIMIC-CDM上平均提升准确率11.2%,部分场景达90.4%
  • 适合需要持续迭代的临床AI研发,支持可审计的成长路径

临床诊断是一个复杂的认知过程,依赖于动态的线索获取和持续的专业知识积累。然而,当前大多数人工智能系统与这一现实脱节,将诊断视为单次回溯预测,缺乏可控的改进机制。我们开发了DxEvolve,一种通过交互式深度临床研究流程实现自我演化的诊断代理。该框架能自主请求检查,并将日益增长的诊疗经验外化为可复用的诊断认知基元。在MIMIC-CDM基准上,DxEvolve相较基础模型平均提升诊断准确率11.2%,在读者研究子集上达到90.4%,接近医生参考水平(88.8%)。在独立外部队列中,相比竞争方法,对源队列覆盖类别提升10.2%准确率,未覆盖类别提升17.1%。通过将经验转化为可管控的学习资产,DxEvolve为临床AI的持续进化提供了可问责的路径。

原文摘要 · Abstract (English)

Clinical diagnosis is a complex cognitive process, grounded in dynamic cue acquisition and continuous expertise accumulation. Yet most current artificial intelligence (AI) systems are misaligned with this reality, treating diagnosis as single-pass retrospective prediction while lacking auditable mechanisms for governed improvement. We developed DxEvolve, a self-evolving diagnostic agent that bridges these gaps through an interactive deep clinical research workflow. The framework autonomously requisitions examinations and continually externalizes clinical experience from increasing encounter exposure as diagnostic cognition primitives. On the MIMIC-CDM benchmark, DxEvolve improved diagnostic accuracy by 11.2% on average over backbone models and reached 90.4% on a reader-study subset, comparable to the clinician reference (88.8%). DxEvolve improved accuracy on an independent external cohort by 10.2% (categories covered by the source cohort) and 17.1% (uncovered categories) compared to the competitive method. By transforming experience into a governable learning asset, DxEvolve supports an accountable pathway for the continual evolution of clinical AI.

AI医疗自我进化诊断优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。