arXiv:2512.11749cs.CV2025-12被引 12

无需VAE,直接在视觉基础模型特征空间生成图像

SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

  • 跳过传统VAE,用视觉基础模型特征直接生成图像
  • 在GenEval上达0.75,在DPG-Bench上达85.78
  • 开源完整训练与生成流程,推动特征驱动图像生成

基于视觉基础模型(VFM)表示的视觉生成,为整合视觉理解、感知与生成提供了一条极具前景的统一路径。尽管潜力巨大,完全在VFM表示空间中训练大规模文本到图像扩散模型仍鲜有探索。为此,我们扩展了SVG(自监督视觉生成表示)框架,提出SVG-T2I,支持在VFM特征域内直接进行高质量文本到图像合成。通过采用标准文本到图像扩散流程,SVG-T2I实现了具有竞争力的表现,在GenEval上达到0.75,在DPG-Bench上达到85.78,验证了VFMs在生成任务中的内在表征能力。我们全面开源该项目,包含自动编码器、生成模型及其训练、推理、评估流程和预训练权重,以促进基于表示的视觉生成研究。

原文摘要 · Abstract (English)

Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training large-scale text-to-image diffusion models entirely within the VFM representation space remains largely unexplored. To bridge this gap, we scale the SVG (Self-supervised representations for Visual Generation) framework, proposing SVG-T2I to support high-quality text-to-image synthesis directly in the VFM feature domain. By leveraging a standard text-to-image diffusion pipeline, SVG-T2I achieves competitive performance, reaching 0.75 on GenEval and 85.78 on DPG-Bench. This performance validates the intrinsic representational power of VFMs for generative tasks. We fully open-source the project, including the autoencoder and generation model, together with their training, inference, evaluation pipelines, and pre-trained weights, to facilitate further research in representation-driven visual generation.

文本生成图像视觉基础模型扩散模型无VAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。