通过调控噪声暴露和网络路径,让扩散自编码器学习更优的潜在表征。
Steering Optimisation Trajectories in Diffusion Representation Learning

- 设计门控残差U-Net与噪声暴露课程,引导优化轨迹。
- 在多个基准上提升表征质量,降低随机种子敏感性。
- 适用于物体中心学习中的空间解耦,改善分割效果。
我们研究了为何扩散自编码器在图像质量相近的情况下仍能学习到显著不同的潜在结构。分析表明,这种现象源于优化动态;我们追踪重建图像与潜在表示质量的关系曲线,发现训练初期存在两种截然不同的优化轨迹。处于重建主导区的模型早期优先保证图像保真度,而处于解耦主导区的模型则更渐进地同时提升重建与解耦能力。我们推测,可通过靶向扩散U-Net中的捷径路径并控制早期噪声水平暴露,来调控这一权衡。为此,我们提出SteeringDRL方法,结合门控残差U-Net与简单噪声暴露课程进行训练。在多个解耦基准测试中,SteeringDRL显著提升了表征质量并降低了对随机种子的敏感性。该方法进一步扩展至物体中心学习中的空间解耦任务,在合成与真实数据集上均提升了分割性能。
原文摘要 · Abstract (English)
We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise around two distinct regimes early in training. Models in the reconstruction regime prioritise image fidelity early, whereas those in the disentanglement regime improve reconstruction and disentanglement more gradually. We hypothesise that this behaviour can be influenced by targeting shortcut pathways in the diffusion U-Net and controlling early noise-level exposure, thereby shaping the reconstruction-disentanglement trade-off during training. To steer optimisation toward stronger representations, we introduce SteeringDRL, combining gated residual U-Nets with a simple noise-level exposure curriculum for training. Across disentanglement benchmarks, SteeringDRL improves representation quality and reduces seed sensitivity. Our method further extends to spatial disentanglement in object-centric learning, improving segmentation quality on synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。