arXiv:2501.10011cs.CVcs.AI2025-01

用多视角3D图像缓解视觉语言模型对物体属性的幻觉

Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions

  • 通过生成3D模型并采样多视角图像作为视觉提示
  • 多视角输入顺序影响模型表现,新方法可消除此影响
  • 引入否定指令降低模型对'是'的偏倚,适合高精度图像理解场景

当前主流的大规模视觉-语言模型(LVLMs)在物体属性判断上存在幻觉问题,导致对输入图像中细粒度属性的错误判定。借助单图生成3D模型的显著进展,本文提出一种新方法以缓解该问题。该方法利用从生成的3D表示中采样的多视角图像作为视觉提示输入给LVLMs,从而提供来自不同视角的更多信息。同时,我们发现多视角图像的输入顺序显著影响模型性能,因此设计了多视角图像增强的视觉语言模型(MIAVLM),其包含一个能够同时消除输入顺序影响并使多视角视觉信息与大语言模型(LLMs)对齐的多视角属性感知模块(MAP)。此外,还设计并使用了否定指令来缓解模型对'是'类回答的偏好。大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Current popular Large Vision-Language Models (LVLMs) are suffering from Hallucinations on Object Attributes (HoOA), leading to incorrect determination of fine-grained attributes in the input images. Leveraging significant advancements in 3D generation from a single image, this paper proposes a novel method to mitigate HoOA in LVLMs. This method utilizes multiview images sampled from generated 3D representations as visual prompts for LVLMs, thereby providing more visual information from other viewpoints. Furthermore, we observe the input order of multiple multiview images significantly affects the performance of LVLMs. Consequently, we have devised Multiview Image Augmented VLM (MIAVLM), incorporating a Multiview Attributes Perceiver (MAP) submodule capable of simultaneously eliminating the influence of input image order and aligning visual information from multiview images with Large Language Models (LLMs). Besides, we designed and employed negative instructions to mitigate LVLMs' bias towards ``Yes" responses. Comprehensive experiments demonstrate the effectiveness of our method.

视觉语言模型幻觉抑制多视角感知3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。