让大模型读懂整张病理切片,提升癌症诊断准确性。
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
- 分三阶段训练,对齐切片与文本、特征空间和任务指令。
- 在18万组问答上测试,形态分析能力显著超越现有模型。
- 适合医学影像分析、病理辅助诊断的研究者使用。
计算病理学近年发展出基于图像块的多模态大模型(MLLM),但这些模型难以全面分析全切片图像(WSI),且常忽略病理科医生依赖的关键形态特征。为此,我们首先提出 WSI-Bench,一个包含来自30种癌症类型的9,850张WSI的大型形态感知基准,涵盖18万组视觉问答对,用于评估模型对诊断关键形态特征的理解能力。在此基础上,我们构建 WSI-LLaVA,一种面向吉比特级全切片图像理解的新框架,采用三阶段训练:切片-文本对齐、特征空间对齐与任务特定指令微调。为更精准评估病理场景下的性能,我们设计了两项专用指标:WSI-Precision 和 WSI-Relevance。实验表明,WSI-LLaVA 在所有能力维度均优于现有模型,尤其在形态分析方面有显著提升,并验证了形态理解与诊断准确性的明确相关性。
原文摘要 · Abstract (English)
Recent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。