arXiv:2412.06774cs.CVcs.AI2024-12CVPR被引 12

将图像细节编码成语言空间的视觉词典,实现高保真图像重建与零样本生成。

Visual Lexicon: Rich Image Features in Language Space

  • 通过自监督学习生成可还原图像的视觉语言标记
  • 单个标记即实现比文本嵌入更高的重建保真度
  • 无需微调即可完成零样本图像生成与视觉理解

我们提出 Visual Lexicon(ViLex),一种新型视觉语言体系,将丰富的图像信息编码至词汇标记的语言空间中,同时保留自然语言难以表达的精细视觉细节。不同于仅关注高层语义(如 CLIP)或像素级重建(如 VAE)的传统方法,ViLex 同时捕捉丰富语义内容与细微视觉特征,支持高质量图像生成与全面的视觉场景理解。通过自监督学习流程,利用冻结的文生图扩散模型优化生成的 ViLex 标记,使其能高效重构输入图像,保留高保真语义重建所需的关键细节。作为语言空间中的图像嵌入,ViLex 标记具备自然语言的组合性,既可独立作为‘文本标记’使用,也可与自然语言标记结合,向预训练文生图模型输入视觉与文本双重信号,模拟与视觉-语言模型交互的方式。实验表明,即使仅用一个 ViLex 标记,其图像重建保真度也高于文本嵌入。此外,ViLex 在无需微调的前提下,以零样本、无监督方式成功完成多种 DreamBooth 任务。同时,作为强大的视觉编码器,其在 15 个基准测试中持续优于强基线 SigLIP 模型。

原文摘要 · Abstract (English)

We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are often challenging to convey in natural language. Unlike traditional methods that prioritize either high-level semantics (e.g., CLIP) or pixel-level reconstruction (e.g., VAE), ViLex simultaneously captures rich semantic content and fine visual details, enabling high-quality image generation and comprehensive visual scene understanding. Through a self-supervised learning pipeline, ViLex generates tokens optimized for reconstructing input images using a frozen text-to-image (T2I) diffusion model, preserving the detailed information necessary for high-fidelity semantic-level reconstruction. As an image embedding in the language space, ViLex tokens leverage the compositionality of natural languages, allowing them to be used independently as "text tokens" or combined with natural language tokens to prompt pretrained T2I models with both visual and textual inputs, mirroring how we interact with vision-language models (VLMs). Experiments demonstrate that ViLex achieves higher fidelity in image reconstruction compared to text embeddings--even with a single ViLex token. Moreover, ViLex successfully performs various DreamBooth tasks in a zero-shot, unsupervised manner without fine-tuning T2I models. Additionally, ViLex serves as a powerful vision encoder, consistently improving vision-language model performance across 15 benchmarks relative to a strong SigLIP baseline.

视觉语言图像编码零样本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。