用扩散模型生成多样视角,提升图像自监督表征学习效果
CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views
- 用图像条件扩散模型生成多样化合成视角作为掩码重建目标
- 在多个基准上超越传统图像裁剪方法,接近视频基方法性能
- 适合缺乏视频数据但需强姿态鲁棒性的自监督学习场景
跨视图掩码自编码已成为学习密集对应关系的有效预训练任务,对视频标签传播等应用至关重要。传统方法以锚定视图重建掩码目标视图为框架,但高质量训练数据获取困难:视频数据采集成本高,简单图像裁剪又缺乏姿态变化,导致性能不足。本文提出CDG-MAE,一种基于MAE的自监督方法,利用图像条件扩散模型从静态图像生成多样合成视图。我们提出量化评估方法,用于选择生成视图在局部与全局上具一致性的扩散模型。生成视图呈现显著的姿态与视角变化,提供丰富训练信号,克服了视频与裁剪方法的局限。此外,我们将标准单锚点掩码策略扩展为多锚点策略,提高预训练任务难度。CDG-MAE显著缩小了与视频基MAE方法的性能差距,同时保持图像仅基方法的数据优势。
原文摘要 · Abstract (English)
Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such as video label propagation. The cross-view pretext task is modeled with a masked autoencoder, where a masked target view is reconstructed from an anchor view. However, acquiring effective training data remains a challenge - collecting diverse video datasets is costly, while simple image crops lack the necessary pose variations, underperforming video-based methods. This paper introduces CDG-MAE, a novel MAE-based self-supervised method that uses diverse synthetic views generated from static images via an image-conditioned diffusion model. We present a quantitative method to evaluate the local and global consistency of the generated views to choose the right diffusion model for cross-view self-supervised pretraining. These generated views exhibit substantial changes in pose and perspective, providing a rich training signal that overcomes the limitations of video and crop-based anchors. Furthermore, we enhance the standard single-anchor MAE setting to a multi-anchor masking strategy to increase the difficulty of the pretext task. CDG-MAE substantially narrows the gap to video-based MAE methods, while maintaining the data advantages of image-only MAEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。