arXiv:2602.17186cs.CV2026-02中稿 · ICML被引 1

用视觉信息增益筛选训练数据,让大模型更依赖图像而非猜答案。

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

  • 基于困惑度设计视觉信息增益指标,量化图像对预测的帮助程度。
  • 仅训练高信息增益样本和词元,性能提升且标注需求减少。
  • 适合需要提升视觉理解、降低语言偏见的多模态研究者。

大型视觉语言模型虽取得显著进展,却常因语言偏见而忽视视觉证据。现有方法通过解码策略、结构修改或精心构建指令数据缓解此问题,但缺乏对每个训练样本或词元实际从图像中获益程度的定量评估。本文提出视觉信息增益(VIG),一种基于困惑度的指标,用于衡量视觉输入带来的预测不确定性降低。VIG可在样本与词元层面进行细粒度分析,有效识别颜色、空间关系、属性等视觉依赖项。基于此,我们提出一种由VIG引导的选择性训练方案,优先处理高VIG样本与词元。该方法显著增强视觉定位能力,缓解语言偏见,在大幅减少监督成本的同时实现更优性能。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying on visual evidence. While prior work attempts to mitigate this issue through decoding strategies, architectural modifications, or curated instruction data, they typically lack a quantitative measure of how much individual training samples or tokens actually benefit from the image. In this work, we introduce Visual Information Gain (VIG), a perplexity-based metric that measures the reduction in prediction uncertainty provided by visual input. VIG enables fine-grained analysis at both sample and token levels, effectively highlighting visually grounded elements such as colors, spatial relations, and attributes. Leveraging this, we propose a VIG-guided selective training scheme that prioritizes high-VIG samples and tokens. This approach improves visual grounding and mitigates language bias, achieving superior performance with significantly reduced supervision by focusing exclusively on visually informative samples and tokens.

视觉语言模型信息增益选择性训练视觉接地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。