用全息表示法让神经网络自动拆解数据中的独立因素,效果优于传统方法。
Disentanglement with Holographic Reduced Representations

- 用全息缩减表示(HRR)构建可微的符号化表示结构
- 在多个指标上达到与主流方法相当的解耦性能
- 适合对可解释性、抗噪性要求高的模型设计场景
解耦学习——即通过神经网络分离数据中不同变化因素——仍是机器学习中的长期难题。以往工作多采用变分自编码器和生成对抗网络,并引入变分推断与信息论约束。本文提出一种新思路:将解耦表示视为符号结构,基于样本概念间的组合关系进行建模。为实现可微学习,我们引入无监督的全息缩减表示(HRR)算法。实验表明,HRR的解绑定操作提供了分离因子的归纳偏置,在潜在空间遍历和解耦度量上表现优异。理论分析进一步证明,解绑定过程能生成近似独立的符号-值对,并推导出每槽容量上限,量化了其促进解耦的归纳偏置。该表示不同于传统自编码器,其潜变量为向量之和而非低维标量。结果表明,该表示对噪声更鲁棒,在多种信噪比下仍保持良好重建质量。
原文摘要 · Abstract (English)
Disentanglement, the separation of factors of variation in data using neural networks, remains a long-standing challenge in machine learning. Prior work has addressed this problem with variational autoencoders and generative adversarial networks that incorporate ideas from variational inference and information-theoretic constraints. In contrast to methods that rely on continuous representations, we propose a design that treats disentangled representations as symbolic structures, motivated by the compositional relationships among the concepts that make up samples from a distribution. However, learning discrete symbolic structures with neural networks while maintaining differentiability is difficult and often requires complex architectures. To address this, we introduce an unsupervised learning algorithm that uses holographic reduced representations (HRR) for neural disentanglement. We show that the HRR unbinding operation provides an inductive bias for separating factors and yields competitive results against baselines, as measured by latent traversals and disentanglement metrics. We complement these empirical findings with an information-theoretic analysis of the HRR unbinding channel. We prove that unbinding induces approximately independent symbol-value pairs and derive a per-slot capacity bound that quantifies how many distinct symbolic concepts can be reliably encoded, giving a quantitative account of the inductive bias toward disentanglement. The resulting representations differ from standard autoencoder-based models, in that their latent units are vectors that are summed together, rather than scalar dimensions of a low-dimensional latent vector. We show that this HRR representation is more robust to noise than other disentangled representations and maintains reconstruction quality across a range of SNRs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。