用分阶段稀疏路由机制,让大模型高效协同专家诊断。
Sparse Multi-Stage Expert-Agent Routing for Complex Clinical Reasoning

- 按阶段动态激活少量专家,根据证据逐步推理。
- 专家调用次数从17次降到3次,诊断准确率达91.5%。
- 适合医疗多学科会诊场景,提升效率与临床可信度。
复杂临床推理需要模型随新证据更新诊断假设,并在资源有限下协调不同专科。现有基于大模型的系统多为单次预测或固定多代理流程,导致专家参与要么静态,要么过度冗余。本文提出稀疏多阶段专家代理路由框架,将诊断建模为分阶段路由过程。随着多模态临床证据逐步提供,框架维护动态病例状态,自适应激活少量专家代理,并通过跨阶段的专家专属记忆支持推理。为超越表面相似性评估自由文本诊断结论,我们引入ClinFEScore——一种面向临床事实的语义评估协议。在重建的多阶段病例(MAC与AgentClinic-NEJM)上,平均激活专家数由17.0降至3.0,同时保持强事实级诊断质量。在200个真实医院多学科会诊病例中,ClinFEScore与医生判断高度相关(Spearman's ρ=0.81;Pearson's r=0.87),本方法实现91.5%的医生验证诊断准确率,每例约5次专家代理/大模型调用。结果表明,稀疏分阶段协调是高效且具临床意义的大模型临床推理方法。
原文摘要 · Abstract (English)
Complex clinical reasoning requires models to update diagnostic hypotheses as new evidence emerges and to coordinate different medical specialities under limited consultation resources. Existing LLM-based clinical reasoning systems typically perform single-pass prediction or rely on fixed multi-agent workflows, making expert participation either static or unnecessarily exhaustive. We propose Sparse Multi-Stage Expert-Agent Routing, a language-based clinical reasoning framework that models diagnosis as a stage-wise routing process. Given progressively available clinical evidence derived from multiple modalities, the framework maintains an evolving case state and adaptively activates a sparse set of medical expert agents, supported by expert-specific memory across stages. To evaluate free-text diagnostic conclusions beyond surface similarity, we further introduce ClinFEScore, a fact-aware semantic evaluation protocol for clinical reasoning outputs. On reconstructed multi-stage cases from MAC and AgentClinic-NEJM, our framework reduces the average number of activated experts from 17.0 to 3.0 whilst maintaining strong fact-level diagnostic quality. On 200 real-world hospital MDT cases, ClinFEScore correlates strongly with clinician judgements (Spearman's $ρ=0.81$; Pearson's $r=0.87$), whilst our method achieves 91.5\% clinician-verified diagnostic accuracy with approximately five expert-agent/LLM calls per case. These results support sparse stage-wise coordination as an efficient and clinically relevant approach to LLM-based clinical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。