arXiv:2606.09653cs.LG2026-06

提出统一框架,揭示概念对齐的四个维度,解决现有方法模糊不清的问题。

A Unifying Framework for Concept-Based Representational Similarity

论文配图:A Unifying Framework for Concept-Based Representational Similarity
图 1 · 摘自论文原文
  • 从表示与概念、实例与分布两个维度重新定义对齐机制
  • 0.1%配对数据即可实现实例级对齐,仅靠无监督目标无法达成
  • 提出CoSAE模型,联合优化多种对齐目标,效果显著提升

不同模型和模态的表征常表现出显著的结构相似性,暗示其共享底层概念分解。然而,概念对齐的定义仍不清晰:现有方法使用相同术语却优化不同目标,导致实际对齐内容模糊。本文提出统一框架,将对齐分解为两个维度:对齐对象(表示 vs. 概念)和对齐层级(实例级 vs. 分布级),由此导出四种属性——实例与分布级别的平移一致性及概念一致性,并揭示现有方法实际保证了哪些性质。我们进一步提出 InterVenchA,一个基于干预的基准,可分别评估概念提取质量、平移质量和概念一致性。理论与实验表明,对齐目标间的常见假设在实践中并不成立:优化某一属性无法可靠恢复其他属性,纯无监督目标无法获得有意义的实例级对齐。为此,我们提出耦合稀疏自编码器(CoSAE),联合施加互补对齐目标,强对齐仅在此场景下出现。令人惊讶的是,仅需0.1%的配对数据即可锚定分布目标并恢复实例级对齐。总体而言,概念对齐本质上是多目标问题,必须作为多目标进行定义、测量与优化。

原文摘要 · Abstract (English)

Learned representations across models and modalities often exhibit striking structural similarities, suggesting shared underlying concept decompositions. However, concept alignment remains poorly defined: existing approaches optimize different objectives under the same terminology, obscuring what is actually aligned. We propose a unifying framework that decomposes alignment along two axes: what is aligned (representations vs. concepts) and at what level (instance-wise vs. distributional). This induces four corresponding properties -- instance-wise and distributional variants of translation and concept consistency -- and reveals precisely which of these guarantees existing methods provide. We further introduce \InterVenchA, an intervention-based benchmark that separately measures extraction quality, translation quality, and concept consistency. Through theory and experiments, we show that commonly assumed equivalences between alignment objectives fail in practice: optimizing one property does not reliably recover the others, and purely unsupervised objectives fail to recover meaningful instance-level alignment. We then propose the Coupled Sparse Autoencoder (CoSAE), which jointly enforces complementary alignment objectives. Strong alignment emerges only in this regime. Surprisingly, as little as 0.1\% paired data is sufficient to recover instance-level alignment when anchoring distributional objectives. Overall, our results show that concept alignment is fundamentally multi-objective: it must be defined, measured, and optimized as such.

表示学习概念对齐多目标优化自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。