arXiv:2502.16223cs.CV2025-02ICLR被引 4

用结构化提示库提升视觉语言模型的零样本医学检测能力

Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection

  • 将提示词分层编码为知识库,实现图文细粒度对齐
  • 在7个基准上比现有方法提升4.1%的平均精度
  • 特别适合无标注医疗图像的零样本检测任务

零样本医学检测可在不依赖标注医学图像的情况下进一步提升检测性能,具有重要临床价值。现有方法利用基于视觉的语言模型(GLIP)通过详细疾病描述作为提示,在推理阶段提升目标疾病名称的识别。然而,这些方法通常将提示视为与目标名称等价的上下文,难以根据视觉信息分配特定疾病知识,导致图像与目标描述之间的对齐粗糙。本文提出StructuralGLIP,引入辅助分支将提示逐层编码为潜在知识库,实现更上下文感知、细粒度的对齐。具体地,在每一层,从图像表示和知识库中选取高度相似特征,构建捕捉图像块与目标描述之间细微关系的结构化表示,并跨模态融合以进一步提升检测性能。大量实验表明,StructuralGLIP在七个零样本医学检测基准上相比现有最佳方法提升4.1%的平均精度(AP),并在内窥镜图像数据集上使微调模型性能持续提升3.2%的AP。

原文摘要 · Abstract (English)

Zero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as prompts for the target disease name during the inference phase. However, these methods typically treat prompts as equivalent context to the target name, making it difficult to assign specific disease knowledge based on visual information, leading to a coarse alignment between images and target descriptions. In this paper, we propose StructuralGLIP, which introduces an auxiliary branch to encode prompts into a latent knowledge bank layer-by-layer, enabling more context-aware and fine-grained alignment. Specifically, in each layer, we select highly similar features from both the image representation and the knowledge bank, forming structural representations that capture nuanced relationships between image patches and target descriptions. These features are then fused across modalities to further enhance detection performance. Extensive experiments demonstrate that StructuralGLIP achieves a +4.1\% AP improvement over prior state-of-the-art methods across seven zero-shot medical detection benchmarks, and consistently improves fine-tuned models by +3.2\% AP on endoscopy image datasets.

零样本检测视觉语言模型医学图像提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。