arXiv:2504.21051cs.LGcs.CL2025-04综述被引 26

综述医学多模态大模型的应用进展与挑战

Multimodal Large Language Models for Medicine: A Comprehensive Survey

  • 系统梳理330篇论文,归纳医疗报告、诊断、治疗三大应用方向
  • 总结六类医疗数据及对应评估基准,揭示模型实际表现
  • 指出临床落地难题并提出可行改进路径,适合研究者参考

多模态大语言模型(MLLMs)近年来成为人工智能研究热点。基于大语言模型的强大能力,MLLMs擅长处理复杂的多模态任务。随着GPT-4的发布,该技术在多个领域引发关注,医学与健康领域亦开始探索其潜力。本文首先介绍LLMs与MLLMs的背景与基本概念,重点阐述其工作原理。随后,系统总结医学领域的三大应用方向:医疗报告生成、医疗诊断与治疗支持。研究基于对330篇近期论文的全面分析,通过具体案例展示MLLMs在这些场景中的卓越能力。针对数据,归纳六种主流数据模式及其对应的评估基准。最后,讨论当前医学领域中MLLMs面临的技术与应用挑战,并提出可行的缓解与解决策略。

原文摘要 · Abstract (English)

MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains. Researchers have begun to explore the potential of MLLMs in the medical and healthcare domain. In this paper, we first introduce the background and fundamental concepts related to LLMs and MLLMs, while emphasizing the working principles of MLLMs. Subsequently, we summarize three main directions of application within healthcare: medical reporting, medical diagnosis, and medical treatment. Our findings are based on a comprehensive review of 330 recent papers in this area. We illustrate the remarkable capabilities of MLLMs in these domains by providing specific examples. For data, we present six mainstream modes of data along with their corresponding evaluation benchmarks. At the end of the survey, we discuss the challenges faced by MLLMs in the medical and healthcare domain and propose feasible methods to mitigate or overcome these issues.

多模态医疗AI大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。