arXiv:2512.04015cs.CV2025-12

自动发现图像隐空间中可变与不变的结构,实现可控变换。

Learning Group Actions In Disentangled Latent Image Representations

  • 用可学习的二值掩码动态划分隐变量,自动识别变化部分
  • 在五个数据集上实现隐空间解耦与群作用联合学习
  • 适用于任意编码器-解码器架构,适合需要可控生成的场景

对隐表示中的群作用进行建模,可实现高维图像数据的可控变换。以往方法多在高维数据空间操作,群作用均匀作用于整个输入,难以解耦受变换影响的子空间。虽然隐空间方法更灵活,但仍需人工划分等变与不变的隐变量,限制了群作用的鲁棒学习。为此,我们提出首个端到端框架,首次在隐图像流形上学习群作用,无需人工干预即可自动发现与变换相关的结构。方法采用可学习的二值掩码结合直通估计,动态划分隐表示为敏感与不变成分,并在统一优化框架中联合学习隐空间解耦与群变换映射。该框架可无缝集成至任意标准编码器-解码器架构。我们在五个2D/3D图像数据集上验证方法,证明其能自动学习多样数据中的解耦隐因子用于群作用,下游分类任务也证实所学表示的有效性。代码已公开于 https://github.com/farhanaswarnali/Learning-Group-Actions-In-Disentangled-Latent-Image-Representations。

原文摘要 · Abstract (English)

Modeling group actions on latent representations enables controllable transformations of high-dimensional image data. Prior works applying group-theoretic priors or modeling transformations typically operate in the high-dimensional data space, where group actions apply uniformly across the entire input, making it difficult to disentangle the subspace that varies under transformations. While latent-space methods offer greater flexibility, they still require manual partitioning of latent variables into equivariant and invariant subspaces, limiting the ability to robustly learn and operate group actions within the representation space. To address this, we introduce a novel end-to-end framework that for the first time learns group actions on latent image manifolds, automatically discovering transformation-relevant structures without manual intervention. Our method uses learnable binary masks with straight-through estimation to dynamically partition latent representations into transformation-sensitive and invariant components. We formulate this within a unified optimization framework that jointly learns latent disentanglement and group transformation mappings. The framework can be seamlessly integrated with any standard encoder-decoder architecture. We validate our approach on five 2D/3D image datasets, demonstrating its ability to automatically learn disentangled latent factors for group actions in diverse data, while downstream classification tasks confirm the effectiveness of the learned representations. Our code is publicly available at https://github.com/farhanaswarnali/Learning-Group-Actions-In-Disentangled-Latent-Image-Representations .

隐空间建模解耦表示群作用可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。