用中间层上下文嵌入提升视觉幻觉检测与定位精度
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs
- 利用LMM中间层的上下文嵌入替代传统对数透镜方法
- 在动作、OCR等多类幻觉上检测准确率显著提升
- 适合需要精准视觉定位和上下文理解的研究者
大型多模态模型(LMMs)通过结合大语言模型(LLMs)的语言能力与模态专用编码器,大幅推动了多模态理解的发展。然而,幻觉问题严重制约了其可靠性与应用。现有检测方法多依赖昂贵训练或外部模型,而近期基于内部特征的方法展现出潜力。本文批判性分析了最先进的无训练方法——对数透镜在处理通用视觉幻觉时的局限性,提出改进方案ContextualLens,利用LMM中层上下文标记嵌入,显著提升在动作、OCR等多样类别上的幻觉检测与定位性能,同时在空间关系、属性比较等需上下文理解的任务中表现优异。该新定位技术可生成高精度边界框,实现从零样本目标分割到有根基的视觉问答的跃迁。本工作为更可靠、可解释的多模态模型铺平道路。
原文摘要 · Abstract (English)
The rapid development of Large Multimodal Models (LMMs) has significantly advanced multimodal understanding by harnessing the language abilities of Large Language Models (LLMs) and integrating modality-specific encoders. However, LMMs are plagued by hallucinations that limit their reliability and adoption. While traditional methods to detect and mitigate these hallucinations often involve costly training or rely heavily on external models, recent approaches utilizing internal model features present a promising alternative. In this paper, we critically assess the limitations of the state-of-the-art training-free technique, the logit lens, in handling generalized visual hallucinations. We introduce ContextualLens, a refined method that leverages contextual token embeddings from middle layers of LMMs. This approach significantly improves hallucination detection and grounding across diverse categories, including actions and OCR, while also excelling in tasks requiring contextual understanding, such as spatial relations and attribute comparison. Our novel grounding technique yields highly precise bounding boxes, facilitating a transition from Zero-Shot Object Segmentation to Grounded Visual Question Answering. Our contributions pave the way for more reliable and interpretable multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。