arXiv:2411.12195cs.CV2024-11IJCV综述被引 21

综述医学视觉语言模型在诊疗中的应用与技术进展。

A Survey of Medical Vision-and-Language Applications and Their Techniques

  • 梳理医学多模态模型的跨模态融合策略
  • 对比不同模型在医疗任务中的性能表现
  • 适合医学AI研究者与临床辅助系统开发者

医学视觉-语言模型(MVLMs)因其能为复杂医学数据提供自然语言接口而备受关注。它们可提升个体患者诊断准确率和决策能力,并通过高效分析大规模数据促进公共卫生监测、疾病追踪及政策制定。MVLMs将自然语言处理与医学图像结合,实现对医学图像及其对应文本信息的综合理解。与通用视觉-语言模型不同,MVLMs专为医学领域设计,能自动提取并解读医学影像与报告中的关键信息以支持临床决策。典型应用场景包括自动报告生成、医学视觉问答、多模态分割、诊断与预后预测、图像-文本检索。本文全面综述了MVLMs及其在各类医疗任务中的应用,详细分析了不同模型架构的跨模态整合策略,评估了常用数据集与标准化指标下的模型表现,并探讨潜在挑战与未来研究方向。相关论文与代码汇总见:https://github.com/YtongXie/Medical-Vision-and-Language-Tasks-and-Methodologies-A-Survey。

原文摘要 · Abstract (English)

Medical vision-and-language models (MVLMs) have attracted substantial interest due to their capability to offer a natural language interface for interpreting complex medical data. Their applications are versatile and have the potential to improve diagnostic accuracy and decision-making for individual patients while also contributing to enhanced public health monitoring, disease surveillance, and policy-making through more efficient analysis of large data sets. MVLMS integrate natural language processing with medical images to enable a more comprehensive and contextual understanding of medical images alongside their corresponding textual information. Unlike general vision-and-language models trained on diverse, non-specialized datasets, MVLMs are purpose-built for the medical domain, automatically extracting and interpreting critical information from medical images and textual reports to support clinical decision-making. Popular clinical applications of MVLMs include automated medical report generation, medical visual question answering, medical multimodal segmentation, diagnosis and prognosis and medical image-text retrieval. Here, we provide a comprehensive overview of MVLMs and the various medical tasks to which they have been applied. We conduct a detailed analysis of various vision-and-language model architectures, focusing on their distinct strategies for cross-modal integration/exploitation of medical visual and textual features. We also examine the datasets used for these tasks and compare the performance of different models based on standardized evaluation metrics. Furthermore, we highlight potential challenges and summarize future research trends and directions. The full collection of papers and codes is available at: https://github.com/YtongXie/Medical-Vision-and-Language-Tasks-and-Methodologies-A-Survey.

医学AI多模态视觉语言综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。