对比视觉与文本提示在分割中的表现,揭示各自优劣。
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
- 构建跨14个数据集的基准,统一评估文本与视觉提示
- 文本提示在常见概念上表现更好,但复杂领域效果差
- 视觉提示平均性能稳定但结果波动大,适合特定场景
提示工程在大型语言模型中已取得显著成效,但在计算机视觉领域的系统性探索仍有限。在语义分割任务中,文本提示可通过开放词汇方法实现任意类别分割,而视觉参考提示则提供直观示例。然而,现有基准多孤立评估不同模态,缺乏同等条件下的直接对比。本文提出 Show or Tell (SoT) 基准,专门用于评估14个涵盖7个不同领域(常见场景、城市、食物、垃圾、部件、工具、土地覆盖)数据集上的文本与视觉提示。我们评估了5种开放词汇方法和4种视觉参考提示方法,并通过置信度融合策略适配后者以支持多类别分割。大量实验表明,开放词汇方法在易用文字描述的常见概念上表现优异,但在工具等复杂领域表现不佳;视觉参考提示方法虽平均表现良好,但结果受输入提示影响大,波动明显。通过全面定量与定性分析,我们明确了两类提示的优缺点,为未来视觉基础模型在分割任务中的研究提供重要参考。
原文摘要 · Abstract (English)
Prompt engineering has shown remarkable success with large language models, yet its systematic exploration in computer vision remains limited. In semantic segmentation, both textual and visual prompts offer distinct advantages: textual prompts through open-vocabulary methods allow segmentation of arbitrary categories, while visual reference prompts provide intuitive reference examples. However, existing benchmarks evaluate these modalities in isolation, without direct comparison under identical conditions. We present Show or Tell (SoT), a novel benchmark specifically designed to evaluate both visual and textual prompts for semantic segmentation across 14 datasets spanning 7 diverse domains (common scenes, urban, food, waste, parts, tools, and land-cover). We evaluate 5 open-vocabulary methods and 4 visual reference prompt approaches, adapting the latter to handle multi-class segmentation through a confidence-based mask merging strategy. Our extensive experiments reveal that open-vocabulary methods excel with common concepts easily described by text but struggle with complex domains like tools, while visual reference prompt methods achieve good average results but exhibit high variability depending on the input prompt. Through comprehensive quantitative and qualitative analysis, we identify the strengths and weaknesses of both prompting modalities, providing valuable insights to guide future research in vision foundation models for segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。