arXiv:2410.11701cs.CLcs.AI2024-10被引 1

用简单指令让AI更依赖图像,减少幻觉。

Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions

  • 通过引导模型优先关注图像,抑制基于内知识的错误生成。
  • 在多个数据集上表现优异,效果媲美甚至超过复杂方法。
  • 无需训练,适配主流闭源与开源多模态模型。

多模态大语言模型(MLLMs)中的幻觉现象严重阻碍了其实际应用。为此,我们提出一种名为Magnifier Prompt(MagPrompt)的简单而有效的方法,通过极简指令缓解MLLMs的幻觉问题。MagPrompt基于两大核心原则设计:(1) 模型应更多关注图像内容;(2) 当图像与模型内部知识冲突时,应优先采纳图像信息。该方法无需训练,适用于开源与闭源模型(如GPT-4o和Gemini-pro),在多个数据集上均表现良好,其效果可媲美或优于更复杂的现有方法(如VCD)。此外,我们的提示设计原则与实验分析为理解多模态幻觉提供了重要洞见。

原文摘要 · Abstract (English)

Hallucinations in multimodal large language models (MLLMs) hinder their practical applications. To address this, we propose a Magnifier Prompt (MagPrompt), a simple yet effective method to tackle hallucinations in MLLMs via extremely simple instructions. MagPrompt is based on the following two key principles, which guide the design of various effective prompts, demonstrating robustness: (1) MLLMs should focus more on the image. (2) When there are conflicts between the image and the model's inner knowledge, MLLMs should prioritize the image. MagPrompt is training-free and can be applied to open-source and closed-source models, such as GPT-4o and Gemini-pro. It performs well across many datasets and its effectiveness is comparable or even better than more complex methods like VCD. Furthermore, our prompt design principles and experimental analyses provide valuable insights into multimodal hallucination.

多模态幻觉抑制提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。