arXiv:2502.14047cs.LGcs.AI2025-02ICLR被引 5

从学习理论角度解析模型表征对齐机制,揭示其与核对齐的数学关联。

Towards a Learning Theory of Representation Alignment

  • 引入度量、概率与谱方法统一表征对齐概念
  • 提出拼接(stitching)分析不同表征间关系,发现其与核对齐相关
  • 为表征对齐提供可学习的理论框架,适合研究模型内部机理者

近年来有观点认为,随着人工智能模型规模和性能提升,其表征会趋于对齐。已有实证分析支持此观点,并推测不同表征可能趋向于共同的现实统计模型。本文从学习理论视角出发,重新梳理并连接基于度量、概率和谱的对齐概念。重点聚焦于拼接(stitching)这一方法,用以理解任务背景下不同表征间的交互。主要贡献在于将拼接特性与底层表征的核对齐(kernel alignment)建立联系。研究结果标志着将表征对齐问题纳入学习理论框架的第一步。

原文摘要 · Abstract (English)

It has recently been argued that AI models' representations are becoming aligned as their scale and performance increase. Empirical analyses have been designed to support this idea and conjecture the possible alignment of different representations toward a shared statistical model of reality. In this paper, we propose a learning-theoretic perspective to representation alignment. First, we review and connect different notions of alignment based on metric, probabilistic, and spectral ideas. Then, we focus on stitching, a particular approach to understanding the interplay between different representations in the context of a task. Our main contribution here is relating properties of stitching to the kernel alignment of the underlying representation. Our results can be seen as a first step toward casting representation alignment as a learning-theoretic problem.

表征对齐学习理论核方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。