提升视觉语言模型长文本推理能力,增强对图像信息的依赖
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
- 不需训练,通过剔除次要文本信息来强化视觉依赖
- 在长上下文任务中显著提升多个视觉语言模型性能
- 适用于需要精准图文理解的复杂推理场景
大型视觉语言模型(LVLMs)在跨模态任务中表现优异,但在长上下文推理中因过度依赖文本信息而表现下降。本研究实证分析发现,随着上下文长度增加,模型对语言的依赖上升,而对视觉信息的依赖减弱。为此,我们提出一种无需训练的上下文剪枝方法,有选择性地移除非关键文本内容,从而增强视觉依赖并减少文本噪声。通过构建长上下文数据集验证,该方法在多种LVLM上均有效。进一步分析表明不同剪枝策略具有鲁棒性,并初步探索了剪枝率与上下文长度间的扩展规律。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) excel in cross-model tasks but experience performance declines in long-context reasoning due to overreliance on textual information and reduced visual dependency. In this study, we empirically analyze LVLMs in long-context reasoning, revealing that increased context length leads to a higher dependence on language at the expense of visual dependency. To address this issue, we propose a novel training-free context pruning method that selectively removes less critical textual information. Our approach enhances visual dependency and reduces textual noise, thereby improving LVLM performance in long-context reasoning. We validate our method by constructing a long-context dataset, demonstrating its effectiveness across various LVLMs. Moreover, further analysis confirms the robustness of different token pruning strategies and preliminary explores scaling laws between pruning rates and context length.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。