arXiv:2511.05565cs.CVcs.AI2025-11

用少量标注样本让视觉语言模型学会显微镜下细胞检测

In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy

  • 通过上下文学习让大模型在少样本下完成细胞定位与分类
  • 六次示例后性能提升趋缓,有推理标记的模型更擅长直接定位
  • 为生物显微图像检测提供可复现的基准测试平台

基础视觉-语言模型在自然图像上表现优异,但在生物医学显微成像中的应用仍不充分。本文研究如何利用上下文学习使先进视觉-语言模型在缺乏大规模标注数据的情况下实现少样本目标检测,这在显微图像中尤为常见。我们构建了Micro-OD基准,包含252张精心筛选的图像,涵盖11种细胞类型,来自四个来源,包括两组实验室专家标注的数据集。系统评估了八种VLM在少样本条件下的表现,并比较了含与不含隐式测试时推理标记的变体。进一步提出一种混合少样本目标检测(FSOD)流程,结合检测头与基于VLM的少样本分类器,显著提升了近期VLM在该基准上的性能。跨数据集观察发现,零样本表现较弱,源于领域差距;但加入少样本支持后检测性能持续提升,六次示例后增益趋于平缓。具有推理标记的模型在端到端定位中更有效,而简化版本更适合对预定位图像块进行分类。结果表明,上下文适应是显微成像中的可行路径,我们的基准为推动开放词汇检测在生物医学成像中的发展提供了可复现的测试平台。

原文摘要 · Abstract (English)

Foundation vision-language models (VLMs) excel on natural images, but their utility for biomedical microscopy remains underexplored. In this paper, we investigate how in-context learning enables state-of-the-art VLMs to perform few-shot object detection when large annotated datasets are unavailable, as is often the case with microscopic images. We introduce the Micro-OD benchmark, a curated collection of 252 images specifically curated for in-context learning, with bounding-box annotations spanning 11 cell types across four sources, including two in-lab expert-annotated sets. We systematically evaluate eight VLMs under few-shot conditions and compare variants with and without implicit test-time reasoning tokens. We further implement a hybrid Few-Shot Object Detection (FSOD) pipeline that combines a detection head with a VLM-based few-shot classifier, which enhances the few-shot performance of recent VLMs on our benchmark. Across datasets, we observe that zero-shot performance is weak due to the domain gap; however, few-shot support consistently improves detection, with marginal gains achieved after six shots. We observe that models with reasoning tokens are more effective for end-to-end localization, whereas simpler variants are more suitable for classifying pre-localized crops. Our results highlight in-context adaptation as a practical path for microscopy, and our benchmark provides a reproducible testbed for advancing open-vocabulary detection in biomedical imaging.

视觉语言模型少样本检测显微成像上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。