arXiv:2608.24570cs.AI2026-08

让大模型像医生一样主动找证据,逐步推理诊断。

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

  • 构建动态诊断环境,让模型主动获取和管理临床证据。
  • 在真实病例上诊断准确率提升,且过程更稳定可靠。
  • 适合医疗AI研发者与需要可解释诊断系统的场景。

临床诊断是一个主动的证据获取过程,医生需收集证据、更新假设,并判断证据是否足够作出诊断。然而,多数基于大语言模型的诊断系统仍采用静态的病例到答案预测,缺乏对证据获取的支持。代理型大模型通过工具使用和中间推理路径提供动态解决方案,但现有系统常未能明确患者证据的暴露、结构化与运行时控制方式。本文提出EviDx,一个以证据感知为核心的主动诊断框架,将患者特异性诊断环境与临床诊断结构、观察者引导的运行时管控相结合。其中,$ℎ$-合成从原始临床案例生成交互环境;结构化组织角色专业化代理、证据工具与演进中的证据状态;管控机制通过跟踪不确定性与证据覆盖率来决定诊断终止。通过三级评估金字塔(执行鲁棒性、推理动态性、诊断结果)验证,EviDx在诊断性能与流程稳定性上均有提升,并揭示了模型能力边界。

原文摘要 · Abstract (English)

Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update competing hypotheses, and decide when the available evidence is sufficient for diagnosis. Yet many medical diagnosis systems built around large language models (LLMs) still formulate diagnosis as static case-to-answer prediction, with limited support for evidence acquisition. Agentic LLMs offer a dynamic alternative through tool use and intermediate diagnostic trajectories, but existing systems often under-specify how patient evidence should be exposed, scaffolded, and controlled at runtime. We introduce EviDx, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness. In EviDx, $\mathcal{E}$-Synthesis constructs interactive environments from raw clinical cases; the scaffold organizes role-specialized agents, evidence tools, and evolving evidence states; and the harness regulates diagnostic termination by tracking uncertainty and evidence coverage. A 3-level evaluation pyramid assesses execution robustness, reasoning dynamics, and diagnostic outcomes. Experiments show that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

医学诊断智能代理证据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。