用双线性自编码器发现语言模型中复杂的隐变量流形
Bilinear autoencoders find interpretable manifolds

- 引入双线性结构建模非线性隐空间,捕捉传统线性方法无法发现的几何特征
- 在语言模型上显著降低重构误差,验证多维流形普遍存在且可被有效捕获
- 支持无监督流形发现,适合研究模型内部表示的可解释性与结构
稀疏自编码器已成为揭示神经网络中可解释隐表示的标准工具。然而,显著概念常跨越当前线性方法难以捕捉的流形,需事后分析。本文采用二次型隐变量弥补这一差距:通过双线性自编码器将激活分解为低秩二次形式,在权重空间线性组合,并支持输入无关的几何分析。这种质的区别挑战了标准线性表示假设。实验与可视化表明,多维几何结构极为普遍,复合隐变量能良好捕捉它们,系统性地改善语言模型的重构误差。此外,具有不同几何先验的自编码器虽字典项不同,却恢复相同的输入子空间。这些模型作为无监督流形发现工具,我们通过Qwen 3.5的交互式在线可视化器予以演示。这迈向了非线性但数学可处理的隐表示,其组合设计即具表达力与可解释性。
原文摘要 · Abstract (English)
Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span manifolds that current linear methods cannot capture without post hoc analysis. This paper uses quadratic latents to close this gap: we implement these with bilinear autoencoders, which decompose activations into low-rank quadratic forms, compose linearly in weight space, and admit input-independent geometric analysis. This qualitative difference in what concepts quadratic latents can detect challenges the standard linear representation hypothesis. Our experiments and visualisations show that multi-dimensional geometries are highly prevalent and that composite latents capture them well, systematically improving reconstruction error in language models. Furthermore, we show that autoencoders with varying geometric priors recover the same input subspace despite their dictionary entries being distinct. Practically, these models serve as an unsupervised tool for manifold discovery, which we demonstrate through an interactive online visualizer for Qwen 3.5. This is a step toward nonlinear but mathematically tractable latent representations whose composition is expressive and interpretable by design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。