提出视觉协同去噪新框架,提升像素级扩散模型的语义对齐能力。
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
- 构建统一框架分离关键设计,明确有效成分
- 在ImageNet-256上少训练轮次超越基线与现有方法
- 适合研究生成模型语义对齐与扩散机制的学者
像素空间扩散近期成为潜空间扩散的有力替代,无需预训练自编码器即可生成高质量图像。然而标准像素空间扩散模型语义监督较弱,未显式建模高层视觉结构。近期表示对齐方法(如REPA)表明预训练视觉特征可显著改善扩散训练,视觉协同去噪成为融合此类特征的有前景方向。但现有方法常混杂多重设计选择,难以判断核心要素。本文提出V-Co,基于统一JiT框架系统研究视觉协同去噪。该受控设置使我们能隔离有效成分:第一,保留特征特异性计算并支持灵活跨流交互,采用全双流架构与结构化无条件预测实现分类器自由引导;第二,需更强语义监督与恰当跨流校准,通过感知漂移混合损失与基于RMS的特征重缩放实现。实验在ImageNet-256上显示,相同模型规模下,V-Co优于基础像素扩散模型与强先验像素扩散方法,且训练轮次更少,为未来表示对齐生成模型提供实用指导。
原文摘要 · Abstract (English)
Pixel-space diffusion has recently re-emerged as a strong alternative to latent diffusion, enabling high-quality generation without pretrained autoencoders. However, standard pixel-space diffusion models receive relatively weak semantic supervision and are not explicitly designed to capture high-level visual structure. Recent representation-alignment methods (e.g., REPA) suggest that pretrained visual features can substantially improve diffusion training, and visual co-denoising has emerged as a promising direction for incorporating such features into the generative process. However, existing co-denoising approaches often entangle multiple design choices, making it unclear which are truly essential. We therefore present V-Co, a systematic study of visual co-denoising in a unified JiT-based framework. This controlled setting allows us to isolate the ingredients that make visual co-denoising effective. Our study reveals two main ingredients. First, co-denoising benefits from preserving feature-specific computation while enabling flexible cross-stream interaction, which leads to a fully dual-stream architecture together with a structurally defined unconditional prediction for classifier-free guidance. Second, it requires both stronger semantic supervision and proper cross-stream calibration, which we realize through a perceptual-drifting hybrid loss and RMS-based feature rescaling. Together, these findings yield a simple recipe for visual co-denoising. Experiments on ImageNet-256 show that, at comparable model sizes, V-Co outperforms the underlying pixel-space diffusion baseline and strong prior pixel-diffusion methods while using fewer training epochs, offering practical guidance for future representation-aligned generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。