arXiv:2508.05492cs.LGcs.AI2025-08被引 8

用多个AI专家协作,让医疗多模态数据更准预测疾病

MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction Modelling

  • 分角色分工:图像、检验等数据转成文字摘要
  • 三阶段协作提升预测准确率,超越现有方法
  • 适合需要融合多种医疗数据的临床研究者

多模态电子健康记录(EHR)数据相比单一模态数据能提供更丰富、互补的患者健康信息。然而,由于数据需求量大,有效整合不同模态数据进行临床预测建模仍具挑战。本文提出新型架构Mixture-of-Multimodal-Agents(MoMA),利用多个大语言模型(LLM)代理实现多模态EHR数据的临床预测任务。MoMA采用专用的LLM代理(“专家代理”)将医学影像、实验室检查等非文本模态转化为结构化文本摘要;这些摘要与临床文本共同输入另一LLM(“聚合代理”),生成统一的多模态摘要;最终由第三个LLM(“预测代理”)基于该摘要输出临床预测结果。在三个真实世界数据集上评估不同模态组合与预测场景下的表现,MoMA在多项任务中优于当前最优方法,展现了更高的准确性与任务适应性。

原文摘要 · Abstract (English)

Multimodal electronic health record (EHR) data provide richer, complementary insights into patient health compared to single-modality data. However, effectively integrating diverse data modalities for clinical prediction modeling remains challenging due to the substantial data requirements. We introduce a novel architecture, Mixture-of-Multimodal-Agents (MoMA), designed to leverage multiple large language model (LLM) agents for clinical prediction tasks using multimodal EHR data. MoMA employs specialized LLM agents ("specialist agents") to convert non-textual modalities, such as medical images and laboratory results, into structured textual summaries. These summaries, together with clinical notes, are combined by another LLM ("aggregator agent") to generate a unified multimodal summary, which is then used by a third LLM ("predictor agent") to produce clinical predictions. Evaluating MoMA on three prediction tasks using real-world datasets with different modality combinations and prediction settings, MoMA outperforms current state-of-the-art methods, highlighting its enhanced accuracy and flexibility across various tasks.

医疗AI多模态大模型临床预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。