arXiv:2605.10286cs.AI2026-05中稿 · the AHLI Conferenc…被引 4

评测大模型代理在多模态医疗预测中的表现,发现单代理优于多代理协作。

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks

论文配图:AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks
图 1 · 摘自论文原文
  • 构建多模态医疗预测的代理评估框架,涵盖文本与影像数据。
  • 单代理在多模态任务中表现更优,且校准性更好。
  • 适合研究医疗智能代理、临床决策支持系统的开发者参考。

构建有效的临床决策支持系统需要融合复杂的异构多模态数据,包括时间序列电子健康记录、医学影像、放射科报告和临床笔记。基于大语言模型(LLM)的代理在多种医疗任务中表现出色,尤其在文本类任务中。由于医疗数据分散于不同医院系统,协作式代理框架为缓解数据共享难题提供了新方向。然而,当前对LLM代理在多模态临床风险预测中的有效性仍缺乏系统评估。本文基于大规模真实世界数据,系统评估了LLM代理在单模态与多模态场景下的性能,并量化了单代理与多代理系统之间的差距。结果表明,单代理框架优于简单的多代理系统,在处理多模态数据方面更具优势,且具备更好的校准能力。这凸显了改进多代理协作机制以更好应对异构输入的迫切需求。通过开源代码与评估框架,本工作为医疗领域智能体系统的发展提供了新基准。

原文摘要 · Abstract (English)

Building effective clinical decision support systems requires the synthesis of complex heterogeneous multimodal data. Such modalities include temporal electronic health records data, medical images, radiology reports, and clinical notes. Large language model (LLM)-based agents have shown impressive performance in various healthcare tasks, especially those involving textual modalities. Considering the fragmentation of healthcare data across hospital systems, collaborative agent frameworks present a promising direction to mitigate data sharing challenges. However, the effectiveness of LLM agents for multimodal clinical risk prediction remains largely unexamined. In this work, we conduct a systematic evaluation of LLM-based agents for clinical prediction tasks using large-scale real-world data. We assess performance in unimodal and multimodal settings and quantify performance gaps between single agent and multi-agent systems. Our findings highlight that single agent frameworks outperform naive multi-agent systems, are better at handling multimodal data, and are better calibrated. This underscores a critical need for improving multi-agent collaboration to better handle heterogeneous inputs. By open-sourcing our code and evaluation framework, this work offers a new benchmark to support future developments relating to agentic systems in healthcare.

医疗AI大模型代理多模态临床预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。