arXiv:2601.19179cs.LG2026-01被引 2

提出新自编码器框架,实现非线性降维中的有序主成分表示。

PCAE: Learning Ordered Representations in Latent Space for Intrinsic Dimension Estimation via Principal Component Autoencoder

  • 通过非均匀方差正则化与等距约束结合,使隐空间保持有序结构。
  • 在非线性数据上成功保留各主成分的独立方差,优于传统方法。
  • 适合需要可解释降维的科研人员,尤其关注主成分顺序的场景。

自编码器常被视为主成分分析(PCA)的非线性扩展。已有研究证明,线性自编码器(LAE)可通过非均匀ℓ₂正则化或调整损失函数,恢复PCA中有序、轴对齐的主成分。然而,在非线性设置下,这些方法无法独立捕捉剩余方差,因受非线性映射干扰。本文提出一种新型自编码器框架,结合非均匀方差正则化与等距约束,自然推广了PCA思想,使模型在保持有序表示和方差保留优势的同时,仍适用于非线性降维任务。

原文摘要 · Abstract (English)

Autoencoders have long been considered a nonlinear extension of Principal Component Analysis (PCA). Prior studies have demonstrated that linear autoencoders (LAEs) can recover the ordered, axis-aligned principal components of PCA by incorporating non-uniform $\ell_2$ regularization or by adjusting the loss function. However, these approaches become insufficient in the nonlinear setting, as the remaining variance cannot be properly captured independently of the nonlinear mapping. In this work, we propose a novel autoencoder framework that integrates non-uniform variance regularization with an isometric constraint. This design serves as a natural generalization of PCA, enabling the model to preserve key advantages, such as ordered representations and variance retention, while remaining effective for nonlinear dimensionality reduction tasks.

降维自编码器主成分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。