arXiv:2508.14063cs.IRcs.AI2025-08被引 3

用多智能体模拟专业脑科推理,显著提升AI诊断准确率。

A Multi-Agent Approach to Neurological Clinical Reasoning

  • 将脑科推理拆解为分析、检索、合成、验证四步,由不同智能体协作完成。
  • 基于LLaMA 3.3-70B的多智能体系统达89.2%准确率,比基础模型高19.7个百分点。
  • 特别在高复杂度题目上表现突出,适合需要深度临床推理的场景。

大语言模型在医学领域展现潜力,但其处理神经科专业化推理的能力仍需系统评估。我们构建了一个包含305道以色列神经科执业资格考试题的基准,按事实深度、临床概念整合与推理复杂度三个维度分类。评估了十种大模型,包括基础模型、检索增强生成(RAG)和一种新型多智能体系统。结果显示性能差异显著:OpenAI-o1在基础模型中表现最佳(准确率90.9%),而专用医学模型表现较差(Meditron-70B仅52.9%)。RAG虽有小幅提升,但在复杂推理题上效果有限。相反,我们的多智能体框架将神经科推理分解为问题分析、知识检索、答案合成与验证等专业化认知功能,实现显著提升,尤其对中等能力模型效果明显。基于LLaMA 3.3-70B的多智能体系统达到89.2%准确率,相比其基础模型(69.5%)大幅提升,高复杂度题目进步尤为显著。该方法将不稳定的亚专科表现转化为统一优异,解决了RAG难以突破的推理瓶颈。我们在独立的MedQA 155例神经科病例集上验证了该方法,结果表明,模拟专业化认知过程的结构化多智能体设计能显著增强复杂医疗推理能力,为挑战性临床场景中的AI辅助提供新方向。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown promise in medical domains, but their ability to handle specialized neurological reasoning requires systematic evaluation. We developed a comprehensive benchmark using 305 questions from Israeli Board Certification Exams in Neurology, classified along three complexity dimensions: factual knowledge depth, clinical concept integration, and reasoning complexity. We evaluated ten LLMs using base models, retrieval-augmented generation (RAG), and a novel multi-agent system. Results showed significant performance variation. OpenAI-o1 achieved the highest base performance (90.9% accuracy), while specialized medical models performed poorly (52.9% for Meditron-70B). RAG provided modest benefits but limited effectiveness on complex reasoning questions. In contrast, our multi-agent framework, decomposing neurological reasoning into specialized cognitive functions including question analysis, knowledge retrieval, answer synthesis, and validation, achieved dramatic improvements, especially for mid-range models. The LLaMA 3.3-70B-based agentic system reached 89.2% accuracy versus 69.5% for its base model, with substantial gains on level 3 complexity questions. The multi-agent approach transformed inconsistent subspecialty performance into uniform excellence, addressing neurological reasoning challenges that persisted with RAG enhancement. We validated our approach using an independent dataset of 155 neurological cases from MedQA. Results confirm that structured multi-agent approaches designed to emulate specialized cognitive processes significantly enhance complex medical reasoning, offering promising directions for AI assistance in challenging clinical contexts.

多智能体神经科临床推理LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。