arXiv:2602.03282cs.CVcs.AI2026-02被引 2

现有视觉表征的全局几何分布无法反映组合结构能力,需引入功能敏感性指标。

Global Geometry Is Not Enough for Vision Representations

  • 用输入-输出雅可比矩阵衡量模型对局部输入变化的响应灵敏度
  • 几何指标与组合绑定能力相关性接近零,但雅可比矩阵高度相关
  • 适合关注模型内部机制、可解释性与组合推理的研究者

表示学习中普遍假设全局分布良好的嵌入能支持鲁棒且泛化的表征。这一观念影响了训练目标与评估协议,将全局几何视为表征能力的代理。然而,全局几何虽能编码元素存在性,却常对组合方式不敏感。我们通过测试多种视觉编码器在组合绑定任务中的表现,发现标准几何统计量与组合绑定能力的相关性近乎为零。相反,基于输入-输出雅可比矩阵的功能敏感性指标能可靠追踪该能力。进一步分析表明,这种差异源于损失函数设计:现有损失显式约束嵌入几何,却未约束局部输入-输出映射。结果表明,全局几何仅刻画表征能力的部分维度,功能敏感性是建模复合结构的关键补充维度。

原文摘要 · Abstract (English)

A common assumption in representation learning is that globally well-distributed embeddings support robust and generalizable representations. This focus has shaped both training objectives and evaluation protocols, implicitly treating global geometry as a proxy for representational competence. While global geometry effectively encodes which elements are present, it is often insensitive to how they are composed. We investigate this limitation by testing the ability of geometric metrics to predict compositional binding across a diverse suite of vision encoders. We find that standard geometry-based statistics exhibit near-zero correlation with compositional binding. In contrast, functional sensitivity, as measured by the input--output Jacobian, reliably tracks this capability. We further provide an analytic account showing that this disparity arises from objective design, as existing losses explicitly constrain embedding geometry but leave the local input--output mapping unconstrained. These results suggest that global embedding geometry captures only a partial view of representational competence and establish functional sensitivity as a critical complementary axis for modeling composite structure.

表示学习组合性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。