无需标注数据,通过继续预训练让医学图像模型精准适配新任务
Effortless Vision-Language Model Specialization in Histopathology without Annotation
- 从现有数据库提取图文对,无标注继续预训练
- 零样本和少样本性能显著提升,大样本时媲美有监督方法
- 通用性强,适合快速部署到新病理分析任务
近年来,视觉语言模型(如 CONCH 和 QuiltNet)在病理学中展现出出色的零样本分类能力。但其通用设计在特定下游任务中表现受限。虽然有监督微调可提升性能,但需人工标注。本文提出一种无需标注的适配方法:通过从已有数据库中提取与领域和任务相关的图像-文本对进行继续预训练。在两个模型(CONCH、QuiltNet)上对三个下游任务的实验表明,该方法显著提升零样本和少样本性能。值得注意的是,随着训练数据量增加,其性能可达到少样本方法水平,且无需人工标注。该方法具备任务无关性、高效性和无标注特性,为病理学视觉语言模型的快速适配提供了新路径。代码已开源。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) in histopathology, such as CONCH and QuiltNet, have demonstrated impressive zero-shot classification capabilities across various tasks. However, their general-purpose design may lead to suboptimal performance in specific downstream applications. While supervised fine-tuning methods address this issue, they require manually labeled samples for adaptation. This paper investigates annotation-free adaptation of VLMs through continued pretraining on domain- and task-relevant image-caption pairs extracted from existing databases. Our experiments on two VLMs, CONCH and QuiltNet, across three downstream tasks reveal that these pairs substantially enhance both zero-shot and few-shot performance. Notably, with larger training sizes, continued pretraining matches the performance of few-shot methods while eliminating manual labeling. Its effectiveness, task-agnostic design, and annotation-free workflow make it a promising pathway for adapting VLMs to new histopathology tasks. Code is available at https://github.com/DeepMicroscopy/Annotation-free-VLM-specialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。