arXiv:2603.00047econ.EMcs.AI2026-03被引 3

揭示安全与能力权衡的几何本质,提出可量化计算的对齐税模型。

What Is the Alignment Tax?

  • 用表示空间中的几何关系定义对齐税率,基于安全方向在能力子空间的投影平方。
  • 推导出安全-能力权衡的帕累托前沿,由安全与能力子空间夹角唯一决定。
  • 发现对齐税可分解为数据结构决定的不可约部分和随模型维度衰减的残差项。

对齐税广受讨论但尚未被形式化刻画。本文在表示空间中提出对齐税的几何理论。在线性表示假设下,将对齐税率定义为安全方向在能力子空间上的投影平方,并推导出由安全与能力子空间主夹角参数化的帕累托前沿。证明该前沿是紧的,并具有递归结构。在能力约束下,安全-安全权衡也遵循相同方程,仅将夹角替换为给定能力方向下安全目标间的偏相关系数。我们推导出一个分解标度律,将对齐税拆分为由数据结构决定的不可约成分,以及随模型维度d以O(m'/d)速度趋近于零的填充残差,并确立了能力保持如何调解或解决安全目标间冲突的条件。

原文摘要 · Abstract (English)

The alignment tax is widely discussed but has not been formally characterized. We provide a geometric theory of the alignment tax in representation space. Under linear representation assumptions, we define the alignment tax rate as the squared projection of the safety direction onto the capability subspace and derive the Pareto frontier governing safety-capability tradeoffs, parameterized by a single quantity of the principal angle between the safety and capability subspaces. We prove this frontier is tight and show it has a recursive structure. safety-safety tradeoffs under capability constraints are governed by the same equation, with the angle replaced by the partial correlation between safety objectives given capability directions. We derive a scaling law decomposing the alignment tax into an irreducible component determined by data structure and a packing residual that vanishes as $O(m'/d)$ with model dimension $d$, and establish conditions under which capability preservation mediates or resolves conflicts between safety objectives.

对齐税几何理论安全-能力权衡标度律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。