arXiv:2410.11761cs.CVcs.AI2024-10CVPR被引 71

首个能理解超高分辨率病理切片的多模态对话系统

SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding

  • 构建首个支持全切片图像的视觉语言对话模型,实现跨场景病理问答
  • 在18个任务中达到顶尖性能,最高准确率达81.17%(TCGA数据集)
  • 开源了4200张切片标注与17.6万组视觉问答数据,适合医学AI研究者使用

尽管多模态大语言模型在计算病理学领域取得进展,但其仍主要聚焦于局部切片分析,缺乏对全切片图像(WSIs)整体上下文的理解。全切片图像的吉字节级规模及缺乏大规模指令数据集,成为发展的主要挑战。本文提出SlideChat,首个可理解吉字节级全切片图像的视觉语言助手,具备出色的多模态对话能力,能响应复杂病理场景下的多样化指令。为支持开发,我们构建了目前最大的全切片图像指令数据集SlideInstruction,包含4.2K张切片描述和176K组视觉问答对。同时提出SlideBench多模态评测基准,涵盖描述生成与视觉问答任务,用于评估模型在显微镜、诊断等临床场景下的表现。相比通用与专用多模态大模型,SlideChat在22项任务中的18项达到最先进水平,例如在SlideBench-VQA(TCGA)上达到81.17%的整体准确率,在SlideBench-VQA(BCNB)上达54.15%。代码、数据与模型已公开。

原文摘要 · Abstract (English)

Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat's capabilities in varied clinical settings such as microscopy, diagnosis. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io.

病理分析多模态模型视觉语言医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。