arXiv:2412.12077cs.CV2024-12CVPR被引 49

首个统一病理切片与全片分析的150亿参数大模型,打破任务割裂

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology

  • 构建统一多模态模型,同时处理切片与全片图像
  • 7项任务42数据集上表现超9成,超越专用模型
  • 适配病理场景的CLIP变体,支持零样本与少样本学习

大型多模态模型(LMM)在病理学领域带来显著进展。以往研究多独立训练切片级与全片级(WSI)模型,导致知识难以融合且模型冗余。本文提出首个150亿参数的统一多模态基础模型CPath-Omni,可整合切片与全片图像分析,涵盖分类、视觉问答、图文生成和视觉指代等多样化任务。大量实验表明,其在42个数据集中的39个上实现最优或持平表现,优于或匹配各任务专用模型。此外,我们开发了专用于病理的基于CLIP的视觉处理器CPath-CLIP,首次将多种视觉模型与大语言模型作为文本编码器融合,构建更强大的CLIP模型,在9个零样本和4个少样本数据集上达当前最优性能。结果证明,CPath-Omni具备统一多任务能力,有望推动病理学基础模型的发展。

原文摘要 · Abstract (English)

The emergence of large multimodal models (LMMs) has brought significant advancements to pathology. Previous research has primarily focused on separately training patch-level and whole-slide image (WSI)-level models, limiting the integration of learned knowledge across patches and WSIs, and resulting in redundant models. In this work, we introduce CPath-Omni, the first 15-billion-parameter LMM designed to unify both patch and WSI level image analysis, consolidating a variety of tasks at both levels, including classification, visual question answering, captioning, and visual referring prompting. Extensive experiments demonstrate that CPath-Omni achieves state-of-the-art (SOTA) performance across seven diverse tasks on 39 out of 42 datasets, outperforming or matching task-specific models trained for individual tasks. Additionally, we develop a specialized pathology CLIP-based visual processor for CPath-Omni, CPath-CLIP, which, for the first time, integrates different vision models and incorporates a large language model as a text encoder to build a more powerful CLIP model, which achieves SOTA performance on nine zero-shot and four few-shot datasets. Our findings highlight CPath-Omni's ability to unify diverse pathology tasks, demonstrating its potential to streamline and advance the field of foundation model in pathology.

病理分析多模态大模型图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。