让诊断智能体学会识别真正影响决策的关键证据。
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

- 通过对比成功与失败病历,找出关键诊断证据
- 用反事实干预验证证据对诊断的实际影响
- 构建证据-诊断-行动关系图,指导缺失证据获取
与静态医疗问答不同,长时程诊断捕捉临床实践的序列特性:证据在多轮交互中逐步获取、整合与评估后才得出最终诊断。然而,现有医生智能体常因未获取或未充分使用关键证据而失败。近期代理方法尝试通过复用历史轨迹或提炼记忆来改进,但其诊断提升受限于历史信息可能包含噪声或偶然内容,且缺乏对哪些证据真正驱动决策的验证。为此,我们提出CDEG,一种基于图的框架,从历史诊断轨迹中学习可复用的决策关键证据。CDEG通过对比同一病例的成功与失败轨迹识别候选证据,利用受控反事实干预验证其诊断影响,并将诊断-证据-行动关系结构化为图。推理时,CDEG追踪患者证据状态,检索相关关系,引导缺失证据获取或被忽略证据的重新评估。在多个领域内与分布外基准测试中,采用多种医生代理基线,CDEG始终提升诊断性能,相比基础代理最高实现11.5%准确率提升。结果表明,可靠长时程诊断需超越轨迹级经验复用,转向对真正塑造临床决策因素的证据级学习。
原文摘要 · Abstract (English)
Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis--evidence--action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。