发现连续训练目标能显著提升图像视频生成的组合泛化能力
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
- 用连续分布目标替代离散损失,改善模型对概念组合的理解
- 在MaskGIT上引入JEPA辅助目标,组合生成准确率提升12.3%
- 适合研究生成模型泛化机制或改进视觉生成效果的研究者
组合泛化是视觉生成模型生成已知概念新组合的能力,但其影响因素尚未完全明确。本文系统研究了多种设计选择对图像和视频生成中组合泛化的影响。通过受控实验,识别出两个关键因素:(i) 训练目标是否作用于离散或连续分布;(ii) 条件输入在训练中是否提供关于构成概念的信息。基于这些发现,我们证明,通过引入基于JEPA的辅助连续目标来松弛MaskGIT的离散损失,可显著提升其在离散模型中的组合生成性能。
原文摘要 · Abstract (English)
Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of how various design choices influence compositional generalization in image and video generation in a positive or negative way. Through controlled experiments, we identify two key factors: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) to what extent conditioning provides information about the constituent concepts during training. Building on these insights, we show that relaxing the MaskGIT discrete loss with an auxiliary continuous JEPA-based objective can improve compositional performance in discrete models like MaskGIT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。