突破视觉语言模型的线性可分极限,提升抽象推理能力
Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- 用线性可分上限诊断模型缺陷,区分感知与推理问题
- 多数模型无法超越原始特征的线性分类性能,存在对齐缺口
- 通过对比学习重构视觉流形,实现非线性决策和性能跃升
推动视觉语言模型(VLMs)发展的关键挑战在于:其在抽象推理任务(如Bongard问题)中的失败究竟是源于感知缺陷,还是顶层推理错误。为解耦这两个因素,我们提出以线性可分上限(LSC)为核心的诊断框架,即在线性分类器下对VLM原始视觉嵌入的性能上限。对当前先进VLMs的应用显示,普遍存在“对齐缺口”——多数模型未能在生成任务中超越其表征的线性可分性。少数突破该上限的模型,通过两种机制实现:进一步优化视觉表征使其更线性可分,或执行非线性决策逻辑。我们证明此瓶颈并非根本限制,而是可解决的视觉对齐问题。所提方法在标准下一词预测基础上引入对比目标,重构视觉流形为更一维线性几何结构,在图像间比较任务上显著提升,并使模型在抽象组合推理任务中大幅超越LSC。
原文摘要 · Abstract (English)
A challenge in advancing Visual-Language Models (VLMs) is determining whether their failures on abstract reasoning tasks, such as Bongard problems, stem from flawed perception or faulty top-down reasoning. To disentangle these factors, we introduce a diagnostic framework centered on the Linear Separability Ceiling (LSC), the performance achievable by a linear classifier on a VLM's raw visual embeddings. Applying this framework to state-of-the-art VLMs, we uncover a pervasive ''alignment gap'', where most models fail to generatively outperform the linear separability of their representations. We find that the few models surpassing this ceiling do so via two mechanisms: by further refining visual representations into a more linearly separable format or by executing non-linear decision logic. We demonstrate that this bottleneck is not a fundamental limitation but a solvable visual alignment issue. Our method augments standard next-token prediction with a contrastive objective to restructure the visual manifold into a more one-dimensionally linear geometry, improving image-to-image comparison and enabling models to significantly surpass the LSC on abstract compositional reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。