arXiv:2411.00588cs.LGcs.AI2024-11ICLR被引 14

提出新模型α-TCVAE,提升生成多样性与表示解耦性。

$α$-TCVAE: On the relationship between Disentanglement and Diversity

  • 基于信息论设计新型总相关下界,优化解耦与表征信息量。
  • 在MPI3D-Real等复杂数据集上生成更多样且保真度高的样本。
  • 适用于生成模型与强化学习下游任务,验证实际有效性。

尽管解耦表示在生成建模和表征学习中展现出潜力,其下游实用性仍存争议。近期研究通过对称性重新定义解耦,强调降低潜在空间维度以增强生成能力。然而从信息论视角看,将复杂属性分配给特定潜在变量可能不可行,限制了解耦表示在复杂数据上的应用。本文提出α-TCVAE,一种使用新型总相关(TC)下界优化的变分自编码器,最大化解耦性与潜在变量的信息量。该下界基于信息论构建,推广了β-VAE下界,并可退化为已知的变分信息瓶颈(VIB)与条件熵瓶颈(CEB)的凸组合。我们进行了量化分析,支持解耦表示能带来更好生成能力和多样性。此外,在表示学习与强化学习领域进行下游任务实验,结果表明α-TCVAE持续学习到比基线更解耦的表示,并生成更丰富的观测,同时保持视觉保真度。尤其在最真实的解耦数据集MPI3D-Real上表现显著提升,证实其在复杂数据上的建模能力。最后,将模型直接用于先进模型基强化学习代理Director,在loconav Ant Maze任务中显著提升性能,验证其下游实用性。

原文摘要 · Abstract (English)

While disentangled representations have shown promise in generative modeling and representation learning, their downstream usefulness remains debated. Recent studies re-defined disentanglement through a formal connection to symmetries, emphasizing the ability to reduce latent domains and consequently enhance generative capabilities. However, from an information theory viewpoint, assigning a complex attribute to a specific latent variable may be infeasible, limiting the applicability of disentangled representations to simple datasets. In this work, we introduce $α$-TCVAE, a variational autoencoder optimized using a novel total correlation (TC) lower bound that maximizes disentanglement and latent variables informativeness. The proposed TC bound is grounded in information theory constructs, generalizes the $β$-VAE lower bound, and can be reduced to a convex combination of the known variational information bottleneck (VIB) and conditional entropy bottleneck (CEB) terms. Moreover, we present quantitative analyses that support the idea that disentangled representations lead to better generative capabilities and diversity. Additionally, we perform downstream task experiments from both representation and RL domains to assess our questions from a broader ML perspective. Our results demonstrate that $α$-TCVAE consistently learns more disentangled representations than baselines and generates more diverse observations without sacrificing visual fidelity. Notably, $α$-TCVAE exhibits marked improvements on MPI3D-Real, the most realistic disentangled dataset in our study, confirming its ability to represent complex datasets when maximizing the informativeness of individual variables. Finally, testing the proposed model off-the-shelf on a state-of-the-art model-based RL agent, Director, significantly shows $α$-TCVAE downstream usefulness on the loconav Ant Maze task.

解耦表征生成模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。