arXiv:2411.07784cs.LGcs.CV2024-11被引 19

提出交互不对称性原则,解释概念如何可解耦并组合。

Interaction Asymmetry: A General Principle for Learning Composable Abstractions

  • 用高阶导数的块对角性定义概念内部互动更复杂
  • 理论证明该原则能实现解耦与组合泛化,支持到n=2
  • 适合研究表征学习、生成模型的科研人员

学习概念的解耦表示并以未见方式重新组合,对泛化到域外情况至关重要。然而,支撑这种解耦和组合泛化的概念内在属性仍不明确。本文提出交互不对称性原则:‘同一概念的部分之间互动比不同概念部分之间的互动更复杂’。通过在生成器映射从概念到观测数据的第(n+1)阶导数上施加块对角性条件来形式化该原则,不同阶的‘复杂性’对应不同的n。理论证明,交互不对称性可同时实现解耦与组合泛化。我们的结果统一了近期关于物体概念学习的理论,这些工作可视为n=0或1的特例。我们给出n=2的结果,扩展了原有框架至更灵活的生成函数,并推测相同证明策略可推广至更大n。实践中,该理论建议自编码器在解码时应惩罚潜在容量及概念间互动。我们提出一种基于Transformer的变分自编码器实现,通过新颖的注意力权重正则项来满足上述条件。在包含物体的合成图像数据集上,该模型表现出与使用更显式物体中心先验的现有模型相当的物体解耦能力。

原文摘要 · Abstract (English)

Learning disentangled representations of concepts and re-composing them in unseen ways is crucial for generalizing to out-of-domain situations. However, the underlying properties of concepts that enable such disentanglement and compositional generalization remain poorly understood. In this work, we propose the principle of interaction asymmetry which states: "Parts of the same concept have more complex interactions than parts of different concepts". We formalize this via block diagonality conditions on the $(n+1)$th order derivatives of the generator mapping concepts to observed data, where different orders of "complexity" correspond to different $n$. Using this formalism, we prove that interaction asymmetry enables both disentanglement and compositional generalization. Our results unify recent theoretical results for learning concepts of objects, which we show are recovered as special cases with $n\!=\!0$ or $1$. We provide results for up to $n\!=\!2$, thus extending these prior works to more flexible generator functions, and conjecture that the same proof strategies generalize to larger $n$. Practically, our theory suggests that, to disentangle concepts, an autoencoder should penalize its latent capacity and the interactions between concepts during decoding. We propose an implementation of these criteria using a flexible Transformer-based VAE, with a novel regularizer on the attention weights of the decoder. On synthetic image datasets consisting of objects, we provide evidence that this model can achieve comparable object disentanglement to existing models that use more explicit object-centric priors.

表征学习解耦表示生成模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。