用反事实推理增强AI诊断,让模型像医生一样思考假设变化的影响。
Improving Clinical Diagnosis with Counterfactual Multi-Agent Reasoning
- 引入反事实病例编辑,模拟症状变化测试诊断可靠性
- 提出概率差距度量,量化每项证据对诊断的支撑强度
- 多轮专家辩论机制提升诊断可解释性,适合医疗AI研发者
临床诊断是复杂的推理过程,医生通过收集证据、提出假设并检验替代解释来完成。在医学训练中,反事实提问(如‘若某症状消失,诊断会如何变化’)被用来强化鉴别诊断能力。随着大语言模型(LLM)用于诊断辅助,其推荐结果的可解释性日益重要。然而,现有模型通常仅基于固定证据推理,未显式检验各发现对竞争性诊断的支持或削弱作用。本文提出一种受临床训练启发的反事实多智能体诊断框架,通过修改临床发现并评估其对诊断的影响,使假设检验显式化且基于证据。我们定义了反事实概率差距(Counterfactual Probability Gap),通过测量修改后置信度变化,量化单项发现对诊断的支撑强度。该信号引导多轮专家讨论,帮助智能体挑战无支持的假设,优化鉴别诊断,并生成更可解释的推理路径。在三个诊断基准和七种LLM上,本方法持续优于提示工程与既有多智能体基线,尤其在复杂模糊病例中提升显著。人工评估显示,该框架生成的推理更具临床实用性、可靠性和连贯性。结果表明,融入反事实验证是构建可信医疗决策支持AI的关键一步。
原文摘要 · Abstract (English)
Clinical diagnosis is a complex reasoning process in which clinicians gather evidence, form hypotheses, and test them against alternative explanations. In medical training, this reasoning is explicitly developed through counterfactual questioning--e.g., asking how a diagnosis would change if a key symptom were absent or altered--to strengthen differential diagnosis skills. As large language model (LLM)-based systems are increasingly used for diagnostic support, ensuring the interpretability of their recommendations becomes critical. However, most existing LLM-based diagnostic agents reason over fixed clinical evidence without explicitly testing how individual findings support or weaken competing diagnoses. In this work, we propose a counterfactual multi-agent diagnostic framework inspired by clinician training that makes hypothesis testing explicit and evidence-grounded. Our framework introduces counterfactual case editing to modify clinical findings and evaluate how these changes affect competing diagnoses. We further define the Counterfactual Probability Gap, a method that quantifies how strongly individual findings support a diagnosis by measuring confidence shifts under these edits. These counterfactual signals guide multi-round specialist discussions, enabling agents to challenge unsupported hypotheses, refine differential diagnoses, and produce more interpretable reasoning trajectories. Across three diagnostic benchmarks and seven LLMs, our method consistently improves diagnostic accuracy over prompting and prior multi-agent baselines, with the largest gains observed in complex and ambiguous cases. Human evaluation further indicates that our framework produces more clinically useful, reliable, and coherent reasoning. These results suggest that incorporating counterfactual evidence verification is an important step toward building reliable AI systems for clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。