arXiv:2503.03987cs.CVcs.AI2025-03被引 11

RetinalGPT用大模型精准分析眼底图像,支持定量诊断与病灶定位。

RetinalGPT: A Retinal Clinical Preference Conversational Assistant Powered by Large Vision-Language Models

  • 基于眼底图像数据集与定制指令微调,提升视觉理解能力。
  • 在8个基准数据集上显著优于通用大模型的疾病诊断性能。
  • 适合眼科临床研究者使用,支持可解释的全流程诊疗分析。

多模态大语言模型(MLLM)在处理图像、视频等非文本数据方面表现出色,已有如LLaVA-Med等通用领域MLLM被应用于医疗场景。然而,这些模型在眼底图像的理解与解读方面仍显不足。医学专家强调疾病检测需依赖定量分析。这凸显了通用与医疗专用MLLM之间的差距:前者虽应用广泛,却缺乏精准诊断所需的专有知识。为此,我们提出RetinalGPT,一个面向临床偏好的眼底图像量化分析对话助手。通过构建大规模眼底图像数据集、开发新型数据流水线,并采用定制化视觉指令微调,显著提升了模型的眼底分析能力与医学知识丰富度。RetinalGPT在8个基准眼底图像数据集上的疾病诊断表现远超通用领域MLLM。此外,该模型支持定量分析与病灶定位,是首次将大模型用于可解释、端到端临床研究框架的尝试。代码已开源于https://github.com/Retinal-Research/RetinalGPT。

原文摘要 · Abstract (English)

Recently, Multimodal Large Language Models (MLLMs) have gained significant attention for their remarkable ability to process and analyze non-textual data, such as images, videos, and audio. Notably, several adaptations of general-domain MLLMs to the medical field have been explored, including LLaVA-Med. However, these medical adaptations remain insufficiently advanced in understanding and interpreting retinal images. In contrast, medical experts emphasize the importance of quantitative analyses for disease detection and interpretation. This underscores a gap between general-domain and medical-domain MLLMs: while general-domain MLLMs excel in broad applications, they lack the specialized knowledge necessary for precise diagnostic and interpretative tasks in the medical field. To address these challenges, we introduce \textit{RetinalGPT}, a multimodal conversational assistant for clinically preferred quantitative analysis of retinal images. Specifically, we achieve this by compiling a large retinal image dataset, developing a novel data pipeline, and employing customized visual instruction tuning to enhance both retinal analysis and enrich medical knowledge. In particular, RetinalGPT outperforms MLLM in the generic domain by a large margin in the diagnosis of retinal diseases in 8 benchmark retinal datasets. Beyond disease diagnosis, RetinalGPT features quantitative analyses and lesion localization, representing a pioneering step in leveraging LLMs for an interpretable and end-to-end clinical research framework. The code is available at https://github.com/Retinal-Research/RetinalGPT

眼底图像多模态模型临床辅助定量分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。