用视觉语言模型零样本分割皮肤肿瘤,无需标注即可自动划出病灶区域。
Zero-shot segmentation of skin tumors in whole-slide images with vision-language foundation models
- 通过文本提示集合与冻结的视觉语言模型,实现全幻灯片图像的自动分割。
- 在两个自建数据集上表现良好,准确识别梭形细胞肿瘤和皮肤转移灶。
- 适合病理医生快速定位病灶,减轻标注负担,提升诊断效率。
由于皮肤肿瘤活检存在广泛的形态差异、重叠的组织学模式以及良恶性病变之间的细微差别,其精确标注面临巨大挑战。视觉-语言基础模型(VLMs)在图文对数据上预训练,学习视觉特征与诊断术语间的联合表示,可实现无像素级标注的零样本定位与分类。然而,现有多数VLM在组织病理学中的应用仍局限于幻灯片级别任务或依赖粗粒度交互提示,难以在千兆像素级全幻灯片图像(WSIs)中生成细粒度分割。本文提出一种零样本视觉-语言分割框架ZEUS,完全自动化,利用类别特异性文本提示集合和冻结的VLM编码器,在全幻灯片图像中生成高分辨率肿瘤掩码。通过将每张幻灯片划分为重叠块,提取视觉嵌入,并计算与文本提示的余弦相似度,最终生成分割掩码。我们在两个自建数据集——原发性梭形细胞肿瘤和皮肤转移灶上验证了其性能,揭示了提示设计、领域偏移及机构间差异对模型的影响。ZEUS显著降低标注负担,为下游诊断流程提供可解释、可扩展的肿瘤边界划分。
原文摘要 · Abstract (English)
Accurate annotation of cutaneous neoplasm biopsies represents a major challenge due to their wide morphological variability, overlapping histological patterns, and the subtle distinctions between benign and malignant lesions. Vision-language foundation models (VLMs), pre-trained on paired image-text corpora, learn joint representations that bridge visual features and diagnostic terminology, enabling zero-shot localization and classification of tissue regions without pixel-level labels. However, most existing VLM applications in histopathology remain limited to slide-level tasks or rely on coarse interactive prompts, and they struggle to produce fine-grained segmentations across gigapixel whole-slide images (WSIs). In this work, we introduce a zero-shot visual-language segmentation pipeline for whole-slide images (ZEUS), a fully automated, zero-shot segmentation framework that leverages class-specific textual prompt ensembles and frozen VLM encoders to generate high-resolution tumor masks in WSIs. By partitioning each WSI into overlapping patches, extracting visual embeddings, and computing cosine similarities against text prompts, we generate a final segmentation mask. We demonstrate competitive performance on two in-house datasets, primary spindle cell neoplasms and cutaneous metastases, highlighting the influence of prompt design, domain shifts, and institutional variability in VLMs for histopathology. ZEUS markedly reduces annotation burden while offering scalable, explainable tumor delineation for downstream diagnostic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。