arXiv:2605.24699cs.AIcs.LG2026-05

MDIA通过多智能体架构提升医疗诊断模型表现,效果优于纯提示工程。

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

论文配图:MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
图 1 · 摘自论文原文
  • 构建7节点专科路由推理图,实现多轮上下文保持与安全用药控制。
  • 在HealthBench Professional上达0.6272准确率,比ChatGPT for Clinicians高3.72个百分点。
  • 适用于需高可靠医疗决策的场景,尤其关注系统架构设计价值。

多数现有临床基准性能提升归因于提示工程,但我们的结果表明,架构与引擎级设计可带来更大改进。我们提出MDIA,一种基于7节点专科路由临床推理图的多智能体诊断智能体,在完整的HealthBench Professional基准(n = 525)上运行于未微调的大语言模型。MDIA在OpenAI GPT-5.4-2026-03-05下达到0.6272的准确率,较OpenAI ChatGPT for Clinicians高出3.72个百分点。实验显示性能提升源于系统架构:专科路由、多轮上下文保留、药物状态安全过滤、站点筛选搜索、长度感知合成及引擎级可靠性。这些发现支持观点:代理型临床基准表现由基础模型与编排架构共同决定。然而,使用不同评分模型时亦出现显著差异;例如使用Gemini 2.5 Pro作为评分器时,MDIA得分为0.6585,表明评分器选择是评估变异的重要来源。因此,鲁棒的LLM评估需跨多个独立评分模型进行。

原文摘要 · Abstract (English)

Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can come from architectural and engine-level design. We present MDIA, a Multi-agent Diagnostic Intelligence Agent implemented as a 7-node specialty-routed clinical reasoning graph, on the full HealthBench Professional benchmark (n = 525), on a non-fine-tuned LLM. MDIA achieves 0.6272 under OpenAI's GPT-5.4-2026-03-05, which is +3.72 pp above the performance of OpenAI's ChatGPT for Clinicians. The experimental work shows that performance lift is attributable to system architecture: specialty routing, multi-turn context preservation, drug-state safety gating, site-filtered search, length-aware synthesis, and engine-level reliability. These findings support the view that agentic clinical benchmark performance is shaped both by the underlying foundation model and the orchestration architecture. Nevertheless, we also noticed notable differences when using other models as a grader; in particular, when using Gemini 2.5 Pro, MDIA scored 0.6585, which suggests that the choice of grader is a source of variability. Robust evaluation of LLMs would therefore require assessment across several independent grader models.

多智能体医疗AI评测架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。