用扩散模型对齐视觉语言模型与图像编码器的局部特征,实现3D生成的高效条件控制。
GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation

- 通过扩散模型将VLM潜空间对齐到图像编码器的完整局部特征空间
- 仅用通用图文对训练,无需大规模3D数据,仍可生成高质量3D资产
- 支持零样本多模态提示,适合希望模块化集成大模型的研究者
近期将视觉语言模型(VLM)作为生成模型条件输入的提示编码器的方法,通常依赖昂贵的端到端训练或映射特征至压缩表示,丢失了3D资产生成等几何感知任务所需的密集空间结构。为此,我们提出GAP3D,一种模块化、基于扩散的方案,直接将VLM生成的潜变量对齐至预训练图像编码器的完整补丁级特征空间,使冻结的下游生成模型能利用VLM作为提示编码器,同时保持空间结构化的条件信号。在3D资产生成上的评估表明,该方法主要基于通用领域图文对进行训练,无需大规模3D数据。它还展现出对多模态提示的涌现零样本能力,尽管训练仅使用文本输入。尽管当前更关注高层语义而非细粒度细节,GAP3D证明了通过基于扩散的对齐,可部分弥合VLM与图像编码器特征空间间的表示鸿沟,为通过生成对齐实现基础模型的模块化整合迈出了第一步。
原文摘要 · Abstract (English)
Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end training or map features to compressed representations, discarding the dense spatial structure required for geometry-aware tasks like 3D asset generation. To address this, we propose GAP3D, a modular, diffusion-based approach that aligns VLM-generated latents directly to the complete, patch-level feature space of a pre-trained image encoder, enabling a frozen downstream generative model to utilize a VLM as prompt encoder while maintaining a spatially structured conditioning signal. Evaluated on 3D asset generation, our method bypasses the need for large-scale 3D data by training mainly on general-domain image-text pairs. It also exhibits emergent zero-shot capabilities for multimodal prompts, despite being trained exclusively on text input. Finally, while currently prioritizing high-level semantics over fine-grained detail, GAP3D demonstrates that the representation gap between VLM and image-encoder feature spaces can be partially bridged through diffusion-based alignment, taking the first steps towards a modular integration of foundation models through generative alignment to dense embedding spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。