arXiv:2505.09139cs.CV2025-05被引 4

用新指标自动选最优提示,提升视觉语言模型的物体识别准确率

Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models

  • 基于对比类别对齐得分筛选提示,避免混淆类干扰
  • 无需训练或标注数据,检测准确率显著提升
  • 适合希望省去人工调参的视觉语言模型使用者

视觉语言模型(VLMs)通过自然语言提示实现灵活的物体检测,但性能受提示表述影响较大。本文提出一种自动化提示优化方法,利用新颖的对比类别对齐得分(CCAS)对提示进行排序,该得分衡量提示与目标类别语义对齐程度,同时惩罚与混淆类别相似的提示。通过大语言模型生成多样提示候选,并使用句子嵌入模型计算的CCAS进行过滤。在具有挑战性的物体类别上评估表明,该方法可自动选择高精度提示,显著提升物体检测准确率,且无需额外模型训练或标注数据。该可扩展、模型无关的流程为视觉语言模型检测系统提供了无需人工调参的可靠替代方案。

原文摘要 · Abstract (English)

Vision-language models (VLMs) offer flexible object detection through natural language prompts but suffer from performance variability depending on prompt phrasing. In this paper, we introduce a method for automated prompt refinement using a novel metric called the Contrastive Class Alignment Score (CCAS), which ranks prompts based on their semantic alignment with a target object class while penalizing similarity to confounding classes. Our method generates diverse prompt candidates via a large language model and filters them through CCAS, computed using prompt embeddings from a sentence transformer. We evaluate our approach on challenging object categories, demonstrating that our automatic selection of high-precision prompts improves object detection accuracy without the need for additional model training or labeled data. This scalable and model-agnostic pipeline offers a principled alternative to manual prompt engineering for VLM-based detection systems.

视觉语言模型提示优化物体检测自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。