自动发现提升大模型视觉理解的图像提示策略
Visual Prompt Discovery via Semantic Exploration
- 用抽象语义空间+智能选代,自动化探索视觉提示
- 在盲测与BLINK基准上准确率显著超越基线方法
- 适合需要提升视觉推理能力的研究者和开发者
大型视觉语言模型(LVLMs)在图像理解与视觉推理中面临重大挑战,常导致关键感知失误。视觉提示通过引入图像操作代码,展现出缓解这些问题的潜力。然而,现有方法多关注工具选择,而非诊断和解决感知失败的根本原因。由于LVLM的黑箱特性和不可预测性,最优视觉提示需通过实验发现,传统依赖人工试错,效率低下。本文提出一种自动化语义探索框架SEVEX,实现任务级视觉提示的高效发现。该方法利用抽象思想空间作为搜索空间,结合新颖性导向选择算法与语义反馈驱动的创意生成过程,有效应对长篇低级代码干扰和庞大无结构提示空间的问题。在BlindTest和BLINK两个评估LVLM感知能力的基准上,SEVEX在任务准确率、推理效率、探索效率与稳定性上均显著优于基线。特别地,框架发现了复杂且反直觉的视觉策略,突破常规工具使用范式,为通过自动化任务级提示增强LVLM感知提供了新路径。
原文摘要 · Abstract (English)
LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts, which incorporate image manipulation code, have shown promising potential in mitigating these issues. While emerged as a promising direction, previous methods for visual prompt generation have focused on tool selection rather than diagnosing and mitigating the root causes of LVLM perception failures. Because of the opacity and unpredictability of LVLMs, optimal visual prompts must be discovered through empirical experiments, which have relied on manual human trial-and-error. We propose an automated semantic exploration framework for discovering task-wise visual prompts. Our approach enables diverse yet efficient exploration through agent-driven experiments, minimizing human intervention and avoiding the inefficiency of per-sample generation. We introduce a semantic exploration algorithm named SEVEX, which addresses two major challenges of visual prompt exploration: (1) the distraction caused by lengthy, low-level code and (2) the vast, unstructured search space of visual prompts. Specifically, our method leverages an abstract idea space as a search space, a novelty-guided selection algorithm, and a semantic feedback-driven ideation process to efficiently explore diverse visual prompts based on empirical results. We evaluate SEVEX on the BlindTest and BLINK benchmarks, which are designed to assess LVLM perception. Experimental results demonstrate that SEVEX significantly outperforms baseline methods in task accuracy, inference efficiency, exploration efficiency, and exploration stability. Notably, our framework discovers sophisticated and counter-intuitive visual strategies that go beyond conventional tool usage, offering a new paradigm for enhancing LVLM perception through automated, task-wise visual prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。