提出新方法降低视觉语言模型幻觉,提升图像理解准确性。
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- 用动态路由网络融合多专家视觉特征,自适应选择最佳信息。
- 在1万样本的细粒度幻觉基准上,显著减少各类幻觉现象。
- 适合关注模型可信度与视觉推理能力的研究者和开发者。
大型视觉语言模型中的物体幻觉严重制约其实际应用。作为准确解析视觉信息的核心组件,视觉编码器的选择至关重要。我们假设,不同视觉编码器采用的多样化训练范式赋予它们不同的归纳偏置,从而导致其幻觉表现各异。现有基准通常只关注粗粒度幻觉检测,难以捕捉我们假设中的多样化幻觉。为系统分析这些影响,我们引入VHBench-10,一个包含约10,000个样本的综合性基准,用于评估大型视觉语言模型在十种细粒度幻觉类别上的表现。评估结果证实,不同编码器展现出独特的幻觉特征。基于这些发现以及简单特征融合的次优性,我们提出VisionWeaver——一种新型上下文感知路由网络。它利用全局视觉特征生成路由信号,动态聚合来自多个专用专家的视觉特征。全面实验验证了VisionWeaver在显著降低幻觉并提升整体模型性能方面的有效性。
原文摘要 · Abstract (English)
Object hallucination in Large Vision-Language Models (LVLMs) significantly impedes their real-world applicability. As the primary component for accurately interpreting visual information, the choice of visual encoder is pivotal. We hypothesize that the diverse training paradigms employed by different visual encoders instill them with distinct inductive biases, which leads to their diverse hallucination performances. Existing benchmarks typically focus on coarse-grained hallucination detection and fail to capture the diverse hallucinations elaborated in our hypothesis. To systematically analyze these effects, we introduce VHBench-10, a comprehensive benchmark with approximately 10,000 samples for evaluating LVLMs across ten fine-grained hallucination categories. Our evaluations confirm encoders exhibit unique hallucination characteristics. Building on these insights and the suboptimality of simple feature fusion, we propose VisionWeaver, a novel Context-Aware Routing Network. It employs global visual features to generate routing signals, dynamically aggregating visual features from multiple specialized experts. Comprehensive experiments confirm the effectiveness of VisionWeaver in significantly reducing hallucinations and improving overall model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。