通过分层结构提升视觉语言模型的精准定位能力
Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding
- 采用全局感知与局部精修两层架构,模拟人类由粗到细的认知过程
- 在GQA和RefCOCO+等数据集上准确率超越Flamingo等主流模型
- 适合需要精确图像区域理解的复杂视觉问答任务
大型语言模型(LLMs)和视觉语言大模型(LVLMs)在自然语言处理和多模态理解方面取得了显著进展。尽管具备出色的泛化能力,当前的LVLM在复杂现实场景中仍存在鲁棒性不足、幻觉频发和推理错误等问题,尤其是在需要精确图像区域定位和细粒度视觉推理时表现不佳。为此,我们提出分层上下文定位视觉语言模型(HCG-LVLM),其架构模仿人类由粗到细的认知过程。该模型包含两层:全局上下文感知层用于初步整体理解,细粒度局部定位层则引入局部细节增强模块提取高分辨率特征,并通过语义一致性验证器确保视觉-语言对齐的准确性。通过自适应融合机制,两层信息被整合以生成稳健且精确的输出。在GQA、A-OKVQA等细粒度视觉问答数据集,以及RefCOCO/+/g等指代表达理解数据集上的大量实验表明,HCG-LVLM持续优于Flamingo、BLIP-2和MiniGPT-4等先进模型,实现了更高准确率并显著降低幻觉现象,验证了其分层设计在提升细粒度视觉语言理解与精准定位能力方面的有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) have achieved remarkable progress in natural language processing and multimodal understanding. Despite their impressive generalization capabilities, current LVLMs often exhibit insufficient robustness, proneness to hallucination, and reasoning errors in complex real-world scenarios, particularly when precise image region localization and fine-grained visual reasoning are required. To address these limitations, we propose the Hierarchical Contextual Grounding LVLM (HCG-LVLM), a novel architecture that mimics human coarse-to-fine cognitive processing. HCG-LVLM employs a two-layered approach: a Global Contextual Perception layer for initial broad understanding and a Fine-grained Local Grounding layer. The latter incorporates a Local Detail Enhancement Module to extract high-resolution features and a Semantic Consistency Validator to ensure accurate, hallucination-free visual-language alignment. Through an adaptive fusion mechanism, information from both layers is integrated for robust and precise outputs. Extensive experiments on challenging datasets, including GQA, A-OKVQA for fine-grained VQA, and RefCOCO/+/g for Referring Expression Comprehension, demonstrate that HCG-LVLM consistently outperforms state-of-the-art models such as Flamingo, BLIP-2, and MiniGPT-4. Our model achieves superior accuracy and significantly reduces hallucination, validating the effectiveness of its hierarchical design in enhancing fine-grained visual-language understanding and precise grounding capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。