arXiv:2411.03755cs.LGcs.AI2024-11ICLR被引 3

无需知道内容与风格维度,也能准确分离多域数据中的内容和风格特征。

Content-Style Learning from Unaligned Domains: Identifiability under Unknown Latent Dimensions

  • 通过跨域潜在分布匹配,放宽了传统方法对独立性和维度已知的苛刻要求。
  • 在施加稀疏约束条件下,即使未知内容与风格的具体维度,仍能实现可识别性。
  • 首次从理论上证明无维度先验的可行性,适合无监督表示学习研究者参考。

理解未对齐多域数据中潜在内容与风格变量的可识别性,对域转换和数据生成等任务至关重要。现有方法通常依赖严格假设,如各潜变量相互独立且内容与风格维度已知。本文提出一种基于跨域潜在分布匹配(LDM)的新分析框架,在更宽松条件下实现内容-风格可识别性。我们证明,可移除变量分量独立性的限制;更重要的是,若对学习到的潜在表示施加适当的稀疏约束,则无需预先知晓内容与风格的维度即可保证可识别性。这一突破长期困扰无监督表示学习领域,本文首次从理论和实践上验证其可行性。在实现层面,我们将LDM重构为带有耦合潜变量的正则化多域GAN损失函数,在温和条件下等价于LDM,同时显著降低计算开销。实验结果验证了理论主张。

原文摘要 · Abstract (English)

Understanding identifiability of latent content and style variables from unaligned multi-domain data is essential for tasks such as domain translation and data generation. Existing works on content-style identification were often developed under somewhat stringent conditions, e.g., that all latent components are mutually independent and that the dimensions of the content and style variables are known. We introduce a new analytical framework via cross-domain \textit{latent distribution matching} (LDM), which establishes content-style identifiability under substantially more relaxed conditions. Specifically, we show that restrictive assumptions such as component-wise independence of the latent variables can be removed. Most notably, we prove that prior knowledge of the content and style dimensions is not necessary for ensuring identifiability, if sparsity constraints are properly imposed onto the learned latent representations. Bypassing the knowledge of the exact latent dimension has been a longstanding aspiration in unsupervised representation learning -- our analysis is the first to underpin its theoretical and practical viability. On the implementation side, we recast the LDM formulation into a regularized multi-domain GAN loss with coupled latent variables. We show that the reformulation is equivalent to LDM under mild conditions -- yet requiring considerably less computational resource. Experiments corroborate with our theoretical claims.

表示学习无监督可识别性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。