arXiv:2506.19133cs.LGq-bio.QM2025-06中稿 · Transactions on Ma…被引 5

用几何优化方法直接学习数据的非欧空间潜变量,无需编码器。

Riemannian Generative Decoder

  • 用黎曼优化器联合训练解码器,直接生成流形上的潜变量。
  • 在三种真实数据中验证,潜变量严格符合预设几何结构。
  • 无需编码器,兼容现有模型,潜空间可解释性强。

欧几里得表示会扭曲具有内在非欧结构的数据。虽然黎曼表示学习通过将数据嵌入匹配流形来解决此问题,但通常依赖编码器估计选定流形上的密度,这涉及数值不稳定的优化目标,可能损害模型训练和质量。为彻底避免此问题,我们提出黎曼生成解码器,一种在任意黎曼流形上寻找流形值潜变量的统一方法。潜变量由黎曼优化器学习,并与解码器网络联合训练。通过摒弃编码器,我们大幅简化了流形约束,相比现有方法仅能处理少数特定流形而言,具有显著优势。我们在三个案例研究中验证该方法:一个合成分支扩散过程、基于线粒体DNA推断的人类迁徙,以及经历细胞分裂周期的细胞。每个案例均表明,所学表示尊重预设几何结构并捕捉内在非欧特性。本方法仅需解码器,兼容现有架构,且生成与数据几何一致的可解释潜空间。代码已公开于 https://github.com/yhsure/riemannian-generative-decoder。

原文摘要 · Abstract (English)

Euclidean representations distort data with intrinsic non-Euclidean structure. While Riemannian representation learning offers a solution by embedding data onto matching manifolds, it typically relies on an encoder to estimate densities on chosen manifolds. This involves optimizing numerically brittle objectives, potentially harming model training and quality. To completely circumvent this issue, we introduce the Riemannian generative decoder, a unifying approach for finding manifold-valued latents on any Riemannian manifold. Latents are learned with a Riemannian optimizer while jointly training a decoder network. By discarding the encoder, we vastly simplify the manifold constraint compared to current approaches which often only handle few specific manifolds. We validate our approach on three case studies -- a synthetic branching diffusion process, human migrations inferred from mitochondrial DNA, and cells undergoing a cell division cycle -- each showing that learned representations respect the prescribed geometry and capture intrinsic non-Euclidean structure. Our method requires only a decoder, is compatible with existing architectures, and yields interpretable latent spaces aligned with data geometry. Code available on https://github.com/yhsure/riemannian-generative-decoder.

生成模型流形学习几何深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。