arXiv:2512.04207cs.AI2025-12

用多智能体系统帮医生快速识别潜在的严重头痛,提升诊断准确性。

Orchestrator Multi-Agent Clinical Decision Support System for Secondary Headache Diagnosis in Primary Care

  • 分七个专业领域智能体协同分析,每步有依据可追溯。
  • 在90例真实病例中,多智能体系统比单模型高出12%的准确率。
  • 适合基层医疗场景,让有限时间内的医生也能做出可靠判断。

与大多数原发性头痛不同,继发性头痛需特殊治疗,若未能及时处理可能造成严重后果。临床指南指出多个‘警示征象’,如雷击样起病、脑膜刺激征、视乳头水肿、局灶性神经功能缺损、颞动脉炎表现、全身性疾病及‘一生中最剧烈的头痛’等。尽管有指南,但初级保健中判断哪些患者需紧急评估仍具挑战。临床医生常面临时间有限、信息不全和症状多样等问题,易导致漏诊或误治。本文提出一种基于大语言模型(LLM)的多智能体临床决策支持系统,采用编排-专精架构,从自由文本病历中实现明确且可解释的继发性头痛诊断。系统将诊断任务分解为七个领域专精智能体,每个生成结构化、证据支撑的推理过程,中央编排器负责任务分解与智能体调度。我们使用90个专家验证的继发性头痛病例进行评估,并对比了单个LLM基线在两种提示策略下的表现:基于问题的提示(QPrompt)与基于临床指南的提示(GPrompt)。测试了五种开源模型(Qwen-30B、GPT-OSS-20B、Qwen-14B、Qwen-8B 和 Llama-3.1-8B),结果表明,采用GPrompt的编排式多智能体系统在所有模型中均取得最高F1分数,尤其在小模型上优势更明显。研究证明,结构化多智能体推理显著提升诊断准确率,超越单纯提示工程,提供透明、贴近临床的可解释决策支持方案。

原文摘要 · Abstract (English)

Unlike most primary headaches, secondary headaches need specialized care and can have devastating consequences if not treated promptly. Clinical guidelines highlight several 'red flag' features, such as thunderclap onset, meningismus, papilledema, focal neurologic deficits, signs of temporal arteritis, systemic illness, and the 'worst headache of their life' presentation. Despite these guidelines, determining which patients require urgent evaluation remains challenging in primary care settings. Clinicians often work with limited time, incomplete information, and diverse symptom presentations, which can lead to under-recognition and inappropriate care. We present a large language model (LLM)-based multi-agent clinical decision support system built on an orchestrator-specialist architecture, designed to perform explicit and interpretable secondary headache diagnosis from free-text clinical vignettes. The multi-agent system decomposes diagnosis into seven domain-specialized agents, each producing a structured and evidence-grounded rationale, while a central orchestrator performs task decomposition and coordinates agent routing. We evaluated the multi-agent system using 90 expert-validated secondary headache cases and compared its performance with a single-LLM baseline across two prompting strategies: question-based prompting (QPrompt) and clinical practice guideline-based prompting (GPrompt). We tested five open-source LLMs (Qwen-30B, GPT-OSS-20B, Qwen-14B, Qwen-8B, and Llama-3.1-8B), and found that the orchestrated multi-agent system with GPrompt consistently achieved the highest F1 scores, with larger gains in smaller models. These findings demonstrate that structured multi-agent reasoning improves accuracy beyond prompt engineering alone and offers a transparent, clinically aligned approach for explainable decision support in secondary headache diagnosis.

临床决策多智能体头痛诊断LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。