通过优化提示设计,提升视觉语言模型在病理诊断中的零样本表现。
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
- 设计系统化提示框架,控制解剖精度与指令形式
- CONCH模型在精确解剖提示下达到最高准确率
- 解剖上下文对病理分析至关重要,提示越精准效果越好
视觉语言模型(VLMs)因其多模态学习能力,在计算病理学中备受关注,可增强对千兆像素全切片图像(WSI)的大数据分析。然而,其对大规模临床数据、任务设定和提示设计的敏感性仍不明确,尤其在诊断准确性方面。本文系统评估了三种先进VLMs(Quilt-Net、Quilt-LLAVA、CONCH)在包含3,507张千兆像素级全切片图像的消化道病理数据集上的表现,涵盖不同组织类型。通过针对癌侵袭性和异型增生状态的消融研究,构建了涵盖领域特异性、解剖精度、指令框架和输出约束的提示工程框架。结果表明,提示设计显著影响模型性能,当提供精确解剖参考时,CONCH模型表现最佳。同时发现,降低解剖精度会导致性能持续下降。此外,模型复杂度并非决定因素,领域对齐与领域特定训练至关重要。本研究为计算病理中的提示工程提供了基础指南,并凸显了使用适当提示提升诊断准确性的潜力。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have gained significant attention in computational pathology due to their multimodal learning capabilities that enhance big-data analytics of giga-pixel whole slide image (WSI). However, their sensitivity to large-scale clinical data, task formulations, and prompt design remains an open question, particularly in terms of diagnostic accuracy. In this paper, we present a systematic investigation and analysis of three state of the art VLMs for histopathology, namely Quilt-Net, Quilt-LLAVA, and CONCH, on an in-house digestive pathology dataset comprising 3,507 WSIs, each in giga-pixel form, across distinct tissue types. Through a structured ablative study on cancer invasiveness and dysplasia status, we develop a comprehensive prompt engineering framework that systematically varies domain specificity, anatomical precision, instructional framing, and output constraints. Our findings demonstrate that prompt engineering significantly impacts model performance, with the CONCH model achieving the highest accuracy when provided with precise anatomical references. Additionally, we identify the critical importance of anatomical context in histopathological image analysis, as performance consistently degraded when reducing anatomical precision. We also show that model complexity alone does not guarantee superior performance, as effective domain alignment and domain-specific training are critical. These results establish foundational guidelines for prompt engineering in computational pathology and highlight the potential of VLMs to enhance diagnostic accuracy when properly instructed with domain-appropriate prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。