arXiv:2507.07574cs.CV2025-07中稿 · TMLR

突破视觉语言模型的线性可分极限,提升抽象推理能力

Beyond the Linear Separability Ceiling: Aligning Representations in VLMs

  • 用线性可分上限诊断模型缺陷,区分感知与推理问题
  • 多数模型无法超越原始特征的线性分类性能,存在对齐缺口
  • 通过对比学习重构视觉流形,实现非线性决策和性能跃升

推动视觉语言模型(VLMs)发展的关键挑战在于:其在抽象推理任务(如Bongard问题)中的失败究竟是源于感知缺陷,还是顶层推理错误。为解耦这两个因素,我们提出以线性可分上限(LSC)为核心的诊断框架,即在线性分类器下对VLM原始视觉嵌入的性能上限。对当前先进VLMs的应用显示,普遍存在“对齐缺口”——多数模型未能在生成任务中超越其表征的线性可分性。少数突破该上限的模型,通过两种机制实现:进一步优化视觉表征使其更线性可分,或执行非线性决策逻辑。我们证明此瓶颈并非根本限制,而是可解决的视觉对齐问题。所提方法在标准下一词预测基础上引入对比目标,重构视觉流形为更一维线性几何结构,在图像间比较任务上显著提升,并使模型在抽象组合推理任务中大幅超越LSC。

原文摘要 · Abstract (English)

A challenge in advancing Visual-Language Models (VLMs) is determining whether their failures on abstract reasoning tasks, such as Bongard problems, stem from flawed perception or faulty top-down reasoning. To disentangle these factors, we introduce a diagnostic framework centered on the Linear Separability Ceiling (LSC), the performance achievable by a linear classifier on a VLM's raw visual embeddings. Applying this framework to state-of-the-art VLMs, we uncover a pervasive ''alignment gap'', where most models fail to generatively outperform the linear separability of their representations. We find that the few models surpassing this ceiling do so via two mechanisms: by further refining visual representations into a more linearly separable format or by executing non-linear decision logic. We demonstrate that this bottleneck is not a fundamental limitation but a solvable visual alignment issue. Our method augments standard next-token prediction with a contrastive objective to restructure the visual manifold into a more one-dimensionally linear geometry, improving image-to-image comparison and enabling models to significantly surpass the LSC on abstract compositional reasoning tasks.

视觉语言模型抽象推理表示对齐对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。