arXiv:2602.24264cs.CVcs.LG2026-02被引 2

揭示视觉模型组合泛化所需的线性正交表示结构

Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models

  • 理论证明组合泛化需特征线性分解且概念间正交
  • 实证发现现代视觉模型具备部分低秩正交因子结构
  • 适合研究模型表征与泛化能力关系的学者参考

组合泛化指在新情境中识别熟悉成分的能力,是智能系统的核心特性。尽管现代模型在海量数据上训练,仍仅覆盖极小部分输入组合空间,引发对表征结构如何支持未见组合泛化的思考。本文形式化了标准训练下的三个理想条件(可分性、可迁移性、稳定性),并证明它们强制几何约束:表征必须线性分解为各概念成分,且不同概念成分之间必须正交。这为线性表征假说提供了理论基础——神经表征中广泛观察到的线性结构,正是组合泛化的必要结果。进一步推导出概念数量与嵌入几何之间的维数边界。在CLIP、SigLIP、DINO等现代视觉模型上的实证表明,其表征呈现部分线性分解,且各概念因子为低秩、近正交,其结构程度与未见组合的泛化性能高度相关。随着模型规模增长,这些条件预测其可能收敛的表征几何。代码已公开于https://github.com/oshapio/necessary-compositionality。

原文摘要 · Abstract (English)

Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although modern models are trained on massive datasets, they still cover only a tiny fraction of the combinatorial space of possible inputs, raising the question of what structure representations must have to support generalization to unseen combinations. We formalize three desiderata for compositional generalization under standard training (divisibility, transferability, stability) and show they impose necessary geometric constraints: representations must decompose linearly into per-concept components, and these components must be orthogonal across concepts. This provides theoretical grounding for the Linear Representation Hypothesis: the linear structure widely observed in neural representations is a necessary consequence of compositional generalization. We further derive dimension bounds linking the number of composable concepts to the embedding geometry. Empirically, we evaluate these predictions across modern vision models (CLIP, SigLIP, DINO) and find that representations exhibit partial linear factorization with low-rank, near-orthogonal per-concept factors, and that the degree of this structure correlates with compositional generalization on unseen combinations. As models continue to scale, these conditions predict the representational geometry they may converge to. Code is available at https://github.com/oshapio/necessary-compositionality.

表征学习组合泛化正交性视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。