arXiv:2504.21559cs.CVcs.AI2025-04NAACL被引 3

用视觉提示框抑制大模型幻觉,提升图像理解可靠性。

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

  • 通过动态选择最优视觉提示框,实现无模型内部信息的优化。
  • 在POPE和CHAIR数据集上显著降低物体幻觉率。
  • 适用于开源与闭源大模型,无需修改模型结构。

大型视觉语言模型(LVLM)常出现物体幻觉,影响其可靠性。我们发现,简单的基于物体的视觉提示——在图像上叠加边界框、圆形等视觉线索——能显著缓解此类幻觉;但不同视觉提示(VPs)效果差异明显。为此,我们提出黑箱视觉提示工程(BBVPE)框架,无需访问模型内部即可识别最优提示。该方法采用候选提示池,并训练一个路由器模型,根据输入图像动态选择最有效的提示。此黑箱方法具有模型无关性,可应用于开源与专有LVLM。在POPE和CHAIR等基准上的评估表明,BBVPE能有效减少物体幻觉。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting -- overlaying visual cues (e.g., bounding box, circle) on images -- can significantly mitigate such hallucination; however, different visual prompts (VPs) vary in effectiveness. To address this, we propose Black-Box Visual Prompt Engineering (BBVPE), a framework to identify optimal VPs that enhance LVLM responses without needing access to model internals. Our approach employs a pool of candidate VPs and trains a router model to dynamically select the most effective VP for a given input image. This black-box approach is model-agnostic, making it applicable to both open-source and proprietary LVLMs. Evaluations on benchmarks such as POPE and CHAIR demonstrate that BBVPE effectively reduces object hallucination.

视觉提示幻觉抑制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。