arXiv:2502.03123cs.LGcs.AI2025-02

直接学习语义差异,让模型更准确地分离不同特征。

Disentanglement in Difference: Directly Learning Semantically Disentangled Representations by Maximizing Inter-Factor Differences

  • 通过对比损失和差异编码器,直接衡量语义差异。
  • 在dSprites和3DShapes上优于主流方法。
  • 适合需要精准解耦语义特征的研究者。

本文提出一种名为“差异中的解耦”(Disentanglement in Difference, DiD)的新方法,旨在解决潜在变量统计独立性与语义解耦目标之间的内在不一致问题。传统方法通过增强潜在变量间的统计独立性来实现解耦,但统计独立并不等同于语义无关,因此该策略未必提升解耦性能。DiD摒弃对统计独立性的依赖,转而直接学习语义差异:设计差异编码器量化语义区别,并引入对比损失促进跨维度的语义区分。该机制使模型能更准确地分离出不同语义因子。在dSprites和3DShapes数据集上的实验表明,DiD在多种解耦评估指标上均超越现有主流方法。

原文摘要 · Abstract (English)

In this study, Disentanglement in Difference(DiD) is proposed to address the inherent inconsistency between the statistical independence of latent variables and the goal of semantic disentanglement in disentanglement representation learning. Conventional disentanglement methods achieve disentanglement representation by improving statistical independence among latent variables. However, the statistical independence of latent variables does not necessarily imply that they are semantically unrelated, thus, improving statistical independence does not always enhance disentanglement performance. To address the above issue, DiD is proposed to directly learn semantic differences rather than the statistical independence of latent variables. In the DiD, a Difference Encoder is designed to measure the semantic differences; a contrastive loss function is established to facilitate inter-dimensional comparison. Both of them allow the model to directly differentiate and disentangle distinct semantic factors, thereby resolving the inconsistency between statistical independence and semantic disentanglement. Experimental results on the dSprites and 3DShapes datasets demonstrate that the proposed DiD outperforms existing mainstream methods across various disentanglement metrics.

解耦表示语义解耦对比学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。