arXiv:2605.28870cs.LGcs.AI2026-05被引 1

揭示了模型表征对齐的线性本质,解释为何不同模型能跨模态对齐。

Representation Alignment Rests on Linear Structure

论文配图:Representation Alignment Rests on Linear Structure
图 1 · 摘自论文原文
  • 表征中的线性结构是对象与属性关系的体现,稀疏编码可增强对齐效果。
  • 中心化和归一化能有效缓解模型间差异,提升跨模型对齐性能。
  • 数据稀缺导致表征噪声,词频越高对齐越强,揭示噪声来源。

我们通过信号、偏差和噪声三元统计框架研究了柏拉图表征假设(PRH)。信号方面:提出线性表征假设(LRH),认为对象与属性的普遍关系在线性表征中被编码;使用稀疏自编码器提取线性特征,发现其跨模态对齐优于稠密表征。偏差方面:模型因架构与训练差异产生隐式偏差,中心化和归一化能一致改善跨模型对齐。噪声方面:有限样本训练导致表征噪声,发现语言模型与文本嵌入模型中词频与对齐程度存在强正相关。综合三者,提出改进的统计模型,进一步解释多种现代AI架构中表征对齐现象。

原文摘要 · Abstract (English)

We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1) Signal:} We propose that Platonic alignment arises from the universal relationship between objects and attributes, which is encoded linearly in representations according to the Linear Representation Hypothesis (LRH). We provide evidence that LRH helps explain PRH by extracting linear object-attribute features with sparse autoencoders and showing that these sparse representations often exhibit stronger cross-modal alignment than their dense counterparts. {2) Bias:} Models have different implicit biases due to the diverse architectures and training procedures used. We show that this difference can be partially mitigated. Centering and normalization consistently improve cross-model alignment. {3) Noise:} Finite-sample training leads to noise in representations. We provide evidence that representational noise is driven by data scarcity by revealing a strong and consistent positive correlation between word frequency and alignment in LLMs and text embedding models. Synthesizing signal, bias, and noise, we propose a statistical model that refines the Linear Representation Hypothesis and explains further phenomena related to the alignment of representations emerging from diverse modern AI architectures.

表征对齐线性结构语言模型噪声分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。