arXiv:2605.12325cs.CV2026-05中稿 · ICML被引 1

用视觉引导优化文本提示,提升无需训练的图像语义分割效果

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

论文配图:VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference
图 1 · 摘自论文原文
  • 通过视觉反馈动态优化文本提示,增强语义表达能力
  • 在多个数据集上提升1.4%-8.4% mIoU,性能领先
  • 推理开销极小,适合部署于资源受限场景

在不依赖训练的前提下实现高效且泛化性强的开放词汇语义分割仍具挑战,主要源于CLIP模型深层的空间偏差。为突破现有方法局限,本文摒弃基于CLIP的范式,转而采用具备空间感知能力的dino.txt框架,以实现更高效、高质量的密集预测。尽管dino.txt具有强空间感知能力,但其文本查询的语义模糊性导致跨模态交互严重失配。为此,本文提出视觉引导提示进化(VIP),通过别名扩展与视觉引导蒸馏机制挖掘有效语义线索,并以显著性感知方式聚合,生成高保真预测。大量实验表明:1)在平均mIoU上超越当前最优方法1.4%-8.4%;2)在多种复杂场景下表现良好;3)推理时间与内存开销几乎无增加。

原文摘要 · Abstract (English)

Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP-based paradigm and harnesses the recent spatially-aware dino$.$txt framework to facilitate more efficient and high-quality dense prediction. While dino$.$txt exhibits robust spatial awareness, we find that the semantic ambiguity of text queries gives rise to severe mismatch within its dense cross-modal interactions. To address this, we introduce Visual-guided Prompt evolution (VIP) to rectify the semantic expressiveness of text queries in dino$.$txt, unleashing its potential for fine-grained object perception. Towards this end, VIP integrates alias expansion with a visual-guided distillation mechanism to mine valuable semantic cues, which are robustly aggregated in a saliency-aware manner to yield a high-fidelity prediction. Extensive evaluations demonstrate that VIP: 1. surpasses the top-leading methods by 1.4%-8.4% average mIoU, 2. generalizes well to diverse challenging domains, and 3. requires marginal inference time and memory overhead.

语义分割视觉语言提示优化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。