arXiv:2606.05957cs.LGstat.ML2026-06

提出死方向概念,打通奇异学习理论与信息几何的桥梁。

Dead Directions: Geometric Singular Learning

论文配图:Dead Directions: Geometric Singular Learning
图 1 · 摘自论文原文
  • 定义死方向:费雪度量退化的方向,由KL散度衰减速率决定。
  • 无需解析解,直接从参数坐标读取奇异几何的三重指标λ、m、ν。
  • 适用于现代深度网络结构,可从单次前反向传播推断模型复杂度。

奇异学习理论与信息几何长期在不同语言中研究同一类参数空间:前者在解析坐标下计算贝叶斯不变量,后者在原始坐标下工作,但依赖非退化假设,而过参数化模型常违反该假设。本文通过一个基本概念——死方向(死方向)建立二者联系:即费雪度量退化的单位向量,等价于解析奇点集的切向量,其KL阶数由KL散度消失速度决定。两种表述指向同一向量;核心突破在于,该KL阶数可直接从原始参数坐标下沿方向的费雪曲率衰减率恢复,无需希罗纳卡分解。光滑纤维上的选择规则将此衰减速率转化为瓦塔纳贝的单方向对实对数正则化阈值(RLCT)的贡献,并将恢复扩展至多分量交叉、多重性m、奇异波动ν(1维方向下通用)、先验-RLCT偏移及加温后验。进一步,将该速率提升至深层网络:多层K-FAC因子分解将每个费雪块表示为激活侧与梯度侧速率的乘积,体现对偶性,对应于现代网络原语(残差流、层归一化、注意力)。商定理将速率推广至梯度流下的规范商Θ/G;SGD满足条件,标准Adam不满足,因而构建了$G$-等变的Adam族预条件器(DDCAdam)以实现。该桥梁提供对奇异几何的参数坐标操控能力,实现按架构的闭式预测,并能仅凭一次检查点的前向与反向传播,读出瓦塔纳贝三重$(λ, m, ν)$,无需后验采样。

原文摘要 · Abstract (English)

Singular learning theory and information geometry have studied the same parameter spaces in mostly separate vocabularies: the former computes Bayesian invariants in resolved coordinates, the latter works in original coordinates under a non-degeneracy assumption that overparameterised models routinely violate. We bridge them through one primitive, the dead direction: a unit vector along which the Fisher metric degenerates, equivalently a tangent to the analytic singular set with a definite KL order, set by how fast the KL divergence vanishes. The two readings name the same vector; our central move shows its KL order is recoverable as the decay rate of the directional Fisher curvature approaching the singularity, in original parameter coordinates and without a Hironaka resolution. A selection rule on smooth fibres translates this rate into Watanabe's single-direction contribution to the real log canonical threshold, and we extend the recovery to multi-component crossings, multiplicity $m$, the singular fluctuation $ν$ (universal in the KL order for 1D directions), prior-RLCT shifts, and tempered posteriors. We then lift this rate to a deep network: a multi-layer K-FAC factorisation writes each Fisher block as a product of activation- and gradient-side rates with a duality between them, instantiated at modern-network primitives (residual streams, layer normalisation, attention). A quotient theorem carries the rate to the gauge quotient $Θ/G$ under gradient flow on a $G$-invariant metric; SGD qualifies, standard Adam does not, and we construct a $G$-equivariant Adam-family preconditioner (DDCAdam) that does. The bridge yields a parameter-coordinate handle on singular geometry, closed-form per-architecture predictions, and a trajectory-rate readout of Watanabe's triple $(λ, m, ν)$ from one checkpoint's forward and backward passes, without posterior sampling.

奇异学习深度网络信息几何参数分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。