对比视频与图像扩散模型的表征能力,发现视频模型更优。
From Image to Video: An Empirical Study of Diffusion Representations
- 用同一架构训练视频和图像生成模型,比较其表征性能。
- 视频扩散模型在分类、识别等任务上普遍优于图像模型。
- 揭示时间信息对表征学习的关键作用,适合研究生成模型表征者。
扩散模型革新了生成建模,实现了图像与视频合成的惊人真实感。这一成功激发了利用其表征进行视觉理解任务的兴趣。尽管已有研究探索图像生成模型的表征潜力,视频扩散模型的视觉理解能力仍鲜有研究。为此,我们系统比较了同一模型架构分别训练用于视频与图像生成的情况,分析其潜在表示在图像分类、动作识别、深度估计和跟踪等下游任务中的表现。结果表明,视频扩散模型始终优于其图像对应模型,但优势程度存在显著差异。我们进一步分析了不同层、不同噪声水平下提取的特征,以及模型规模和训练预算对表征与生成质量的影响。本工作首次直接比较了视频与图像扩散目标在视觉理解中的表现,揭示了时间信息在表征学习中的关键作用。
原文摘要 · Abstract (English)
Diffusion models have revolutionized generative modeling, enabling unprecedented realism in image and video synthesis. This success has sparked interest in leveraging their representations for visual understanding tasks. While recent works have explored this potential for image generation, the visual understanding capabilities of video diffusion models remain largely uncharted. To address this gap, we systematically compare the same model architecture trained for video versus image generation, analyzing the performance of their latent representations on various downstream tasks including image classification, action recognition, depth estimation, and tracking. Results show that video diffusion models consistently outperform their image counterparts, though we find a striking range in the extent of this superiority. We further analyze features extracted from different layers and with varying noise levels, as well as the effect of model size and training budget on representation and generation quality. This work marks the first direct comparison of video and image diffusion objectives for visual understanding, offering insights into the role of temporal information in representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。