用对抗判别器引导扩散模型分解并重组数据,提升生成质量和可解释性。
Unsupervised Decomposition and Recombination with Discriminator-Driven Diffusion Models
- 通过判别器区分单源与跨源重组样本,驱动生成器学习物理语义一致的成分。
- 在CelebA-HQ等数据集上FID更低,MIG和MCC指标更优,表明解耦更好。
- 首次应用于机器人视频轨迹,重组动作组件显著扩大探索状态空间。
将复杂数据分解为可复用的因子表示,有助于发现通用组件并实现新样本合成。本文研究基于扩散模型的无监督因子分解,在不依赖因子级标注的情况下学习解耦潜在空间。图像中因子可捕捉背景、光照与物体属性;机器人视频中则可提取可复用运动成分。为提升因子发现效果与组合生成质量,我们引入对抗训练信号:训练判别器区分单源样本与跨源重组样本,优化生成器以欺骗判别器,从而促进重组结果在物理与语义上的一致性。方法在CelebA-HQ、Virtual KITTI、CLEVR和Falcor3D上优于基线,获得更低的FID分数,并在MIG和MCC指标上表现更佳。进一步展示新应用:在LIBERO基准上,通过重组学习到的动作组件生成多样化轨迹,显著扩展状态空间覆盖范围,提升探索效率。
原文摘要 · Abstract (English)
Decomposing complex data into factorized representations can reveal reusable components and enable synthesizing new samples via component recombination. We investigate this in the context of diffusion-based models that learn factorized latent spaces without factor-level supervision. In images, factors can capture background, illumination, and object attributes; in robotic videos, they can capture reusable motion components. To improve both latent factor discovery and quality of compositional generation, we introduce an adversarial training signal via a discriminator trained to distinguish between single-source samples and those generated by recombining factors across sources. By optimizing the generator to fool this discriminator, we encourage physical and semantic consistency in the resulting recombinations. Our method outperforms implementations of prior baselines on CelebA-HQ, Virtual KITTI, CLEVR, and Falcor3D, achieving lower FID scores and better disentanglement as measured by MIG and MCC. Furthermore, we demonstrate a novel application to robotic video trajectories: by recombining learned action components, we generate diverse sequences that significantly increase state-space coverage for exploration on the LIBERO benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。