综述生成式AI在医疗中的多模态应用与挑战
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- 系统梳理从单模态到多模态AI的演进路径
- 144篇文献揭示诊断支持与报告生成等关键应用
- 适合关注AI医疗落地的研究者与临床开发者
生成式人工智能(如扩散模型、ChatGPT)正通过提升诊断准确率和自动化临床流程重塑医疗领域。该领域已从仅处理文本的大型语言模型(如临床记录与决策支持),发展为能整合影像、文本、结构化数据等多模态信息的统一系统。本综述遵循PRISMA-ScR指南,系统检索了PubMed、IEEE Xplore与Web of Science,筛选出截至2024年底发表的144篇相关研究。结果表明,多模态方法正成为主流,推动诊断辅助、医学报告生成、药物发现及对话式AI等创新。但数据异构融合、模型可解释性、伦理问题及真实临床环境验证仍面临挑战。本文总结当前技术进展,识别关键空白,为构建可扩展、可信且具临床影响力的多模态AI提供指引。
原文摘要 · Abstract (English)
Generative artificial intelligence (AI) models, such as diffusion models and OpenAI's ChatGPT, are transforming medicine by enhancing diagnostic accuracy and automating clinical workflows. The field has advanced rapidly, evolving from text-only large language models for tasks such as clinical documentation and decision support to multimodal AI systems capable of integrating diverse data modalities, including imaging, text, and structured data, within a single model. The diverse landscape of these technologies, along with rising interest, highlights the need for a comprehensive review of their applications and potential. This scoping review explores the evolution of multimodal AI, highlighting its methods, applications, datasets, and evaluation in clinical settings. Adhering to PRISMA-ScR guidelines, we systematically queried PubMed, IEEE Xplore, and Web of Science, prioritizing recent studies published up to the end of 2024. After rigorous screening, 144 papers were included, revealing key trends and challenges in this dynamic field. Our findings underscore a shift from unimodal to multimodal approaches, driving innovations in diagnostic support, medical report generation, drug discovery, and conversational AI. However, critical challenges remain, including the integration of heterogeneous data types, improving model interpretability, addressing ethical concerns, and validating AI systems in real-world clinical settings. This review summarizes the current state of the art, identifies critical gaps, and provides insights to guide the development of scalable, trustworthy, and clinically impactful multimodal AI solutions in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。