arXiv:2603.01131cs.MAcs.AI2026-03

用分层疾病关系链和多智能体协作提升临床诊断准确性

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

论文配图:MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis
图 1 · 摘自论文原文
  • 基于IBIS协议组织多智能体协同,逐轮推理并支持证据链
  • 在ClinicalBench和MIMIC-IV上诊断准确率优于主流LLM与基线模型
  • 适合需要可解释、高一致性医疗报告的临床辅助场景

临床诊断是逐步整合证据的过程,医生从症状、病史推进到检查、鉴别诊断、疾病关联及治疗决策。尽管大语言模型提升了医学文本理解能力,但其临床应用受限于证据弱关联、推理不透明以及鉴别诊断、最终诊断、诊断依据与治疗计划之间的一致性不足。我们提出MedCollab,一个支持全流程临床诊断与报告生成的多智能体框架。该框架根据患者记录协调专科与检查类智能体,采用问题驱动的信息系统(IBIS)协议组织推理过程,确保每项诊断结论均有患者特异性证据和医学知识支撑。同时构建分层疾病关系链(HDRC),通过进展、并发症与共病关系连接被接受的假设。在多轮辩论中,验证器引导的共识模块评估证据支持度、医学合理性与逻辑矛盾,动态调整智能体贡献并过滤无效推理。在ClinicalBench和MIMIC-IV数据集上的实验表明,MedCollab在诊断准确率、证据一致性和临床推理质量方面均优于领先的LLM与医疗多智能体基线。结果表明,结构化且可审计的协作能生成更真实、更具临床连贯性的诊断报告。

原文摘要 · Abstract (English)

Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations, competing hypotheses, disease relations, and treatment decisions. Large language models have advanced medical text understanding and generation. Yet their clinical use remains limited by weak evidence grounding, opaque reasoning, and inconsistent links among differential diagnosis, final diagnosis, diagnostic basis, and treatment planning. We introduce MedCollab, a multi-agent framework for full-cycle clinical diagnosis and report generation. MedCollab coordinates specialist and examination agents according to patient records. It structures agent deliberation with an Issue-Based Information System (IBIS) protocol, so that each diagnostic position is supported by patient-specific evidence and medical knowledge. It also builds Hierarchical Disease Relation Chains (HDRC) to connect accepted hypotheses through progression, complication, and comorbidity relations. During multi-round deliberation, a verifier-guided consensus module evaluates evidence support, medical plausibility, and logical conflicts. It then adjusts agent contributions and filters unsupported reasoning. Experiments on ClinicalBench and MIMIC-IV show that MedCollab outperforms leading LLMs and medical multi-agent baselines in diagnostic accuracy, evidence consistency, and clinical reasoning quality. These results indicate that structured and auditable collaboration can produce more faithful and clinically coherent diagnostic reports.

多智能体临床诊断疾病关系链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。