arXiv:2501.02669cs.CVcs.CL2025-01ICML被引 13

提升视觉模型推理能力,缓解图像与文本间的认知不平衡问题。

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

  • 设计图像与文本双版本的合成任务,评估视觉语言模型的算法推理能力。
  • 通过图像转文本训练,显著提升模型在复杂任务上的泛化性能。
  • 发现思维链和梯度对齐是促进跨难度迁移的关键机制。

视觉语言模型(VLMs)在图像描述和视觉问答中表现优异,但在多步视觉推理任务上仍远逊于仅用文本处理相同任务的大型语言模型,暴露出模态不平衡或脆弱性问题。为系统研究该现象,我们提出一个合成评估框架,包含三个任务:表格读取、网格导航和视觉类比,每个任务均有 SIMPLE 和 HARD 两个难度层级,即使 SIMPLE 版本也对前沿 VLMs 构成挑战。我们探索了基于 SIMPLE 任务训练策略对 HARD 任务泛化效果的影响,即简单到困难(S2H)泛化。该设置提供等效纯文本版本,可量化模态不平衡及训练策略的影响。实验表明:1)显式图像转文本转换对推动图像任务中的 S2H 泛化至关重要,能将文本中的推理能力迁移至视觉域;2)该转换可在测试时内部完成。此外,我们进行了机制分析,识别出梯度对齐度量可有效预测促进 S2H 泛化的训练策略。消融实验强调了思维链(chain-of-thought)的重要性。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning -- even compared to LLMs on the same tasks presented in text form -- giving rise to perceptions of modality imbalance or brittleness. Towards a systematic study of such issues, we introduce a synthetic framework for assessing the ability of VLMs to perform algorithmic visual reasoning, comprising three tasks: Table Readout, Grid Navigation, and Visual Analogy. Each has two levels of difficulty, SIMPLE and HARD, and even the SIMPLE versions are difficult for frontier VLMs. We propose strategies for training on the SIMPLE version of tasks that improve performance on the corresponding HARD task, i.e., simple-to-hard (S2H) generalization. This controlled setup, where each task also has an equivalent text-only version, allows a quantification of the modality imbalance and how it is impacted by training strategy. We show that 1) explicit image-to-text conversion is important in promoting S2H generalization on images, by transferring reasoning from text; 2) conversion can be internalized at test time. We also report results of mechanistic study of this phenomenon. We identify measures of gradient alignment that can identify training strategies that promote better S2H generalization. Ablations highlight the importance of chain-of-thought.

视觉推理模态平衡图像生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。