arXiv:2412.07767cs.CV2024-12CVPR被引 8

不用文本也能训练出强大的图像生成模型,效果还更好。

Learning Visual Generative Priors without Text

  • 用纯视觉自监督方法训练图像生成模型,无需文本对
  • 仅用十分之一的图文对微调,性能媲美甚至超越现有文本生成模型
  • 适合做图像到3D、视频等无文本依赖任务的前置模型

尽管文本到图像(T2I)模型近年来作为视觉生成先验蓬勃发展,但其对高质量图文配对数据的依赖导致扩展成本高昂。我们认为,构建稳健的视觉生成先验并不需要跨模态对齐,重点应放在纹理建模上。这一理念促使我们研究图像到图像(I2I)生成,使模型能以自监督方式从真实场景图像中学习。我们首先提出纯视觉训练框架Lumos,验证了I2I模型在可行性与可扩展性上的优势。进一步发现,作为T2I的上游任务,我们的I2I模型作为更基础的视觉先验,在仅使用1/10的图文对进行微调时,性能达到或超过现有T2I模型。我们还在图像到3D、图像到视频等与文本无关的任务上,验证了I2I先验优于T2I先验。项目页面见https://ant-research.github.io/lumos。

原文摘要 · Abstract (English)

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that grasping the cross-modality alignment is not a necessity for a sound visual generative prior, whose focus should be on texture modeling. Such a philosophy inspires us to study image-to-image (I2I) generation, where models can learn from in-the-wild images in a self-supervised manner. We first develop a pure vision-based training framework, Lumos, and confirm the feasibility and the scalability of learning I2I models. We then find that, as an upstream task of T2I, our I2I model serves as a more foundational visual prior and achieves on-par or better performance than existing T2I models using only 1/10 text-image pairs for fine-tuning. We further demonstrate the superiority of I2I priors over T2I priors on some text-irrelevant visual generative tasks, like image-to-3D and image-to-video. Our project page is available at https://ant-research.github.io/lumos.

图像生成自监督视觉先验无文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。