系统评估视觉语言模型在检测与分割任务中的表现,揭示其优劣与适用场景。
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
- 首次系统评测多种视觉语言模型在8类检测和8类分割任务中的表现。
- 发现不同微调策略在零预测、视觉微调和文本提示下对性能影响显著。
- 为未来多模态模型设计提供数据支持,适合计算机视觉研究者参考。
视觉语言模型(VLM)在开放词汇(OV)目标检测与分割任务中已广泛应用。尽管在OV任务中展现潜力,其在传统视觉任务中的有效性尚未被系统评估。本文首次将VLM视为基础模型,开展跨多个下游任务的全面评测:1)涵盖8种检测场景(如闭集检测、域自适应、密集物体等)和8种分割场景(如少样本、开放世界、小物体等),揭示不同VLM架构在各类任务中的性能优势与局限;2)针对检测任务,评估三种微调粒度:零预测、视觉微调和文本提示,并分析不同策略在各类任务下的表现差异;3)基于实证结果,深入分析任务特性、模型架构与训练方法间的关联,为未来VLM设计提供洞见;4)本工作对计算机视觉、多模态学习及视觉基础模型领域的研究者具有重要价值,有助于理解当前进展并明确未来方向。相关项目已发布于https://github.com/better-chao/perceptual_abilities_evaluation。
原文摘要 · Abstract (English)
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, their effectiveness in conventional vision tasks has thus far been unevaluated. In this work, we present the systematic review of VLM-based detection and segmentation, view VLM as the foundational model and conduct comprehensive evaluations across multiple downstream tasks for the first time: 1) The evaluation spans eight detection scenarios (closed-set detection, domain adaptation, crowded objects, etc.) and eight segmentation scenarios (few-shot, open-world, small object, etc.), revealing distinct performance advantages and limitations of various VLM architectures across tasks. 2) As for detection tasks, we evaluate VLMs under three finetuning granularities: \textit{zero prediction}, \textit{visual fine-tuning}, and \textit{text prompt}, and further analyze how different finetuning strategies impact performance under varied task. 3) Based on empirical findings, we provide in-depth analysis of the correlations between task characteristics, model architectures, and training methodologies, offering insights for future VLM design. 4) We believe that this work shall be valuable to the pattern recognition experts working in the fields of computer vision, multimodal learning, and vision foundation models by introducing them to the problem, and familiarizing them with the current status of the progress while providing promising directions for future research. A project associated with this review and evaluation has been created at https://github.com/better-chao/perceptual_abilities_evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。