arXiv:2605.23449cs.LGcs.CV2026-05

提出新VAE框架,让模型自动识别并适配隐空间的非对易性结构。

Commutator-Induced Uncertainty in VAEs

论文配图:Commutator-Induced Uncertainty in VAEs
图 1 · 摘自论文原文
  • 用李群几何与代数结合方式诊断隐变量非对易性
  • 通过重构顺序交换测试发现隐空间与重建行为存在尺度不匹配
  • 引入数据驱动的稳定性约束,使重建更符合非对易结构,适合图像生成任务

变分自编码器(VAEs)在学习隐空间的非对易结构时常遇困难。现有对称性感知的VAE通过代数正则化强制对易性,虽适用于对易变换群,但在数据本身具有内在非对易结构时会压制有效信息。本文主张应显式诊断并反映非对易性。提出一种李群变分自编码器框架,融合几何与代数视角的不确定性分析,并分离离散生成因子与连续几何变换。第一阶段在无结构约束下训练,通过有限的Baker-Campbell-Hausdorff偏差和重构顺序交换测试测量代数非对易性与解码器敏感性,发现隐空间非对易性与重建行为间存在尺度不匹配。第二阶段引入数据驱动的变形-稳定性约束,校准解码器敏感度以匹配代数非对易性。在dSprites、3DShapes、3DCars和CelebA上评估,相比通用及对称性感知基线(如beta-VAE、CLG-VAE、CFASL),本方法提升重建质量,解码器行为更符合隐空间非对易结构。定性分析显示更清晰的顺序依赖隐空间组合与更稳定的重构,在CelebA上生成更逼真图像并实现有意义的隐方向交互。

原文摘要 · Abstract (English)

Variational autoencoders (VAEs) often struggle to represent non-commutative structure in learned latent spaces. Symmetry-aware VAEs commonly address this issue by enforcing commutativity through algebraic regularization, which is appropriate for commutative transformation groups but can suppress meaningful non-commutative structure when it is intrinsic to the data. We argue that non-commutativity should instead be explicitly diagnosed and reflected in reconstruction behavior. We introduce a Lie Group VAE framework that combines geometric and algebraic perspectives on uncertainty while separating discrete generative factors from continuous geometric transformations. In a first phase, the model is trained without structural constraints while algebraic non-commutativity is measured through finite Baker-Campbell-Hausdorff deviations and decoder order sensitivity is measured through reconstruction order-swap tests. These diagnostics reveal a scale mismatch between latent non-commutativity and reconstruction behavior under unconstrained training. In a second phase, we introduce a deformation-stability constraint with a data-driven calibration constant that aligns decoder sensitivity with algebraic non-commutativity. We evaluate the framework on dSprites, 3DShapes, 3DCars, and CelebA against generic and symmetry-aware baselines, including beta-VAE, CLG-VAE, and CFASL. Across synthetic benchmarks, the method improves reconstruction quality and yields decoder-level behavior more consistent with latent non-commutative structure. Qualitative analyses show clearer order-dependent latent compositions and more stable reconstructions. On CelebA, the model yields more faithful reconstructions and factor-specific latent traversals than CFASL, while also exhibiting meaningful order-dependent interactions between learned latent directions.

VAE非对易性生成模型隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。