用视觉语言模型提升文本生成3D的语义与空间一致性
Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- 将大视觉语言模型融入扩散采样,实现可微分的语义与空间引导
- 在GPTeval3D上显著提升语义准确率与几何一致性,尤其在多物体场景中
- 适合需要精细控制和复杂关系推理的3D生成任务
Score Distillation Sampling (SDS) 通过多视角2D渲染的去噪过程,利用预训练的文本到图像扩散模型监督3D模型,实现高质量的文本到3D生成,并确保3D一致性。然而现有基于SDS的方法存在两大根本缺陷:(1) 依赖CLIP类文本编码器导致语义对齐粗糙,难以处理细粒度提示;(2) 2D扩散先验缺乏显式3D空间约束,造成几何不一致和多物体场景中的关系错误。为此,我们提出VLM3D,一种新型文本到3D生成框架,将大型视觉语言模型(VLMs)作为可微分的语义与空间先验集成至SDS流程。相较于标准文本到图像扩散先验,VLMs具备丰富的语言-视觉协同监督,实现细粒度提示对齐;其固有的视觉语言建模能力提供强空间理解,显著提升单物体生成的3D一致性,并增强多物体场景中的关系推理能力。我们在开源模型Qwen2.5-VL基础上实现VLM3D,评估其在GPTeval3D基准上的表现。实验表明,在多种物体与复杂场景下,VLM3D在语义保真度、几何连贯性和空间正确性方面均显著优于先前的SDS方法。
原文摘要 · Abstract (English)
Score Distillation Sampling (SDS) enables high-quality text-to-3D generation by supervising 3D models through the denoising of multi-view 2D renderings, using a pretrained text-to-image diffusion model to align with the input prompt and ensure 3D consistency. However, existing SDS-based methods face two fundamental limitations: (1) their reliance on CLIP-style text encoders leads to coarse semantic alignment and struggles with fine-grained prompts; and (2) 2D diffusion priors lack explicit 3D spatial constraints, resulting in geometric inconsistencies and inaccurate object relationships in multi-object scenes. To address these challenges, we propose VLM3D, a novel text-to-3D generation framework that integrates large vision-language models (VLMs) into the SDS pipeline as differentiable semantic and spatial priors. Unlike standard text-to-image diffusion priors, VLMs leverage rich language-grounded supervision that enables fine-grained prompt alignment. Moreover, their inherent vision language modeling provides strong spatial understanding, which significantly enhances 3D consistency for single-object generation and improves relational reasoning in multi-object scenes. We instantiate VLM3D based on the open-source Qwen2.5-VL model and evaluate it on the GPTeval3D benchmark. Experiments across diverse objects and complex scenes show that VLM3D significantly outperforms prior SDS-based methods in semantic fidelity, geometric coherence, and spatial correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。