MedOrch用多个专业工具和推理代理,让AI更灵活地辅助医生诊断。
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
- 构建模块化代理架构,灵活接入医疗专用工具
- 阿尔茨海默病诊断准确率达93.26%,比现有方法高4个百分点
- 支持多模态数据推理,适合复杂临床决策场景
医疗决策是人工智能最具挑战的领域之一,需融合多样知识、复杂推理及外部分析工具。当前AI系统或依赖特定任务模型(适应性差),或使用通用语言模型但缺乏专业知识与工具支撑。本文提出MedOrch框架,通过协调多个专业工具与推理代理,实现全面的医疗决策支持。该框架采用模块化代理架构,可灵活集成领域专用工具而不改变核心系统;同时保证推理过程透明可追溯,便于临床医生逐项验证推荐依据。我们在三种不同医疗任务上评估MedOrch:阿尔茨海默病诊断、胸部X光解读与医学视觉问答,使用真实临床数据集。结果表明,其在各项任务中表现优异:阿尔茨海默病诊断准确率93.26%,超过当前最优基线超4个百分点;疾病进展预测准确率50.35%,显著提升;胸部X光分析宏AUC达61.2%,宏F1得分为25.5%;在复杂多模态视觉问答(图像+表格)任务中,准确率达到54.47%。这些成果凸显MedOrch在驱动工具化推理、处理多模态医疗数据、支持临床认知任务方面的潜力。
原文摘要 · Abstract (English)
Healthcare decision-making represents one of the most challenging domains for Artificial Intelligence (AI), requiring the integration of diverse knowledge sources, complex reasoning, and various external analytical tools. Current AI systems often rely on either task-specific models, which offer limited adaptability, or general language models without grounding with specialized external knowledge and tools. We introduce MedOrch, a novel framework that orchestrates multiple specialized tools and reasoning agents to provide comprehensive medical decision support. MedOrch employs a modular, agent-based architecture that facilitates the flexible integration of domain-specific tools without altering the core system. Furthermore, it ensures transparent and traceable reasoning processes, enabling clinicians to meticulously verify each intermediate step underlying the system's recommendations. We evaluate MedOrch across three distinct medical applications: Alzheimer's disease diagnosis, chest X-ray interpretation, and medical visual question answering, using authentic clinical datasets. The results demonstrate MedOrch's competitive performance across these diverse medical tasks. Notably, in Alzheimer's disease diagnosis, MedOrch achieves an accuracy of 93.26%, surpassing the state-of-the-art baseline by over four percentage points. For predicting Alzheimer's disease progression, it attains a 50.35% accuracy, marking a significant improvement. In chest X-ray analysis, MedOrch exhibits superior performance with a Macro AUC of 61.2% and a Macro F1-score of 25.5%. Moreover, in complex multimodal visual question answering (Image+Table), MedOrch achieves an accuracy of 54.47%. These findings underscore MedOrch's potential to advance healthcare AI by enabling reasoning-driven tool utilization for multimodal medical data processing and supporting intricate cognitive tasks in clinical decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。