arXiv:2601.21568cs.LG2026-01被引 1

用可用信息统一量化表征与功能相似性,揭示双向分析与模型能力的关键作用。

Bridging Functional and Representational Similarity via Usable Information

  • 通过可用信息构建统一框架,连接功能与表征相似性
  • 发现功能比较需双向分析,且相似性依赖预测能力上限
  • 表征相似是功能相似的充分条件,但非必要,适用于模型对比研究

我们提出一个统一框架,通过 extit{可用}信息量化表征间的相似性,实现三个关键维度的严谨理论与实证融合。首先,在功能相似性方面,建立拼接性能与条件互信息之间的形式关联,并揭示拼接本质上具有方向性,强调功能比较需双向分析而非单向映射。其次,在表征相似性方面,发现基于重建的度量方法与标准工具(如CKA、RSA)在特定约束下可作为可用信息的估计器;关键发现是相似性取决于预测族的能力:对严格观察者不同的表征,对更强大的观察者可能等价。第三,证明表征相似性足以保证功能相似性,但非必要。我们通过任务粒度层级统一这些概念:复杂任务上的相似性必然蕴含其粗粒度衍生任务上的相似性,确立输入重建为最大粒度极限下的表征相似性边界。

原文摘要 · Abstract (English)

We present a unified framework for quantifying the similarity between representations through the lens of \textit{usable} information, offering a rigorous theoretical and empirical synthesis across three key dimensions. First, addressing functional similarity, we establish a formal link between stitching performance and conditional mutual information. We further reveal that stitching is inherently asymmetric, demonstrating that robust functional comparison necessitates a bidirectional analysis rather than a unidirectional mapping. Second, concerning representational similarity, we find that reconstruction-based metrics and standard tools (e.g., CKA, RSA) act as estimators of usable information under specific constraints. Crucially, we show that similarity is relative to the capacity of the predictive family: representations that appear distinct to a rigid observer may be identical to a more expressive one. Third, we demonstrate that representational similarity is sufficient but not necessary for functional similarity. We unify these concepts through a task-granularity hierarchy: similarity on a complex task guarantees similarity on any coarser derivative, establishing representational similarity as the limit of maximum granularity: input reconstruction.

表征学习信息论相似性度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。