arXiv:2504.05174cs.LGhep-ph2025-04被引 1

VAE能自动发现数据中的对称性,并压缩到更少的隐变量空间。

Learning symmetries in datasets

  • 通过相关性分析识别隐空间中关键方向,揭示对称性结构
  • 在具有$O(2)$对称性的二维数据上,隐空间维度显著降低
  • 无需监督即可发现物理系统中的隐藏对称性,适合高维数据探索

我们研究了数据集中存在的对称性如何影响变分自编码器(VAE)学习的隐空间结构。在来自简单机械系统和粒子碰撞的数据上训练VAE后,通过相关性度量分析隐空间组织,发现当存在对称性或近似对称性时,VAE会自我组织隐空间,有效沿较少的隐变量压缩数据。这一行为捕捉了由对称性约束决定的内在维度,并揭示特征间的隐藏关系。此外,我们对一个简化玩具模型进行理论分析,证明在理想条件下,隐空间会与数据流形的对称方向对齐。实验涵盖从具有$O(2)$对称性的二维数据到电子-正电子和质子-质子碰撞的真实数据集。结果表明,无监督生成模型可揭示数据中的潜在结构,为无需标注的对称性发现提供新方法。

原文摘要 · Abstract (English)

We investigate how symmetries present in datasets affect the structure of the latent space learned by Variational Autoencoders (VAEs). By training VAEs on data originating from simple mechanical systems and particle collisions, we analyze the organization of the latent space through a relevance measure that identifies the most meaningful latent directions. We show that when symmetries or approximate symmetries are present, the VAE self-organizes its latent space, effectively compressing the data along a reduced number of latent variables. This behavior captures the intrinsic dimensionality determined by the symmetry constraints and reveals hidden relations among the features. Furthermore, we provide a theoretical analysis of a simple toy model, demonstrating how, under idealized conditions, the latent space aligns with the symmetry directions of the data manifold. We illustrate these findings with examples ranging from two-dimensional datasets with $O(2)$ symmetry to realistic datasets from electron-positron and proton-proton collisions. Our results highlight the potential of unsupervised generative models to expose underlying structures in data and offer a novel approach to symmetry discovery without explicit supervision.

隐空间结构对称性发现变分自编码器无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。