arXiv:2511.14271cs.CV2025-11被引 1

用语言模型当质检员,让3D生成更准更合理

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation

论文配图:Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
图 1 · 摘自论文原文
  • 用视觉语言模型生成语义与空间双重评判信号
  • 在标准测试中显著优于现有方法,纠正严重几何错误
  • 适用于优化型和前馈型两种主流3D生成流程

文本到3D生成发展迅速,但当前主流模型(包括基于优化和前馈架构)仍存在两大根本缺陷:一是粗粒度语义对齐困难,难以捕捉提示中的细粒度细节;二是缺乏稳健的3D空间理解,导致部件组装和空间关系出现几何不一致甚至灾难性失败。为此,我们提出VLM3D,一个通用框架,将大型视觉语言模型(VLMs)重新用作可微分的语义与空间评判器。核心贡献是基于VLM的“是/否”对数几率生成双查询评判信号,同时评估语义一致性和几何合理性。我们在两种不同范式中验证了该引导信号的通用性:(1) 作为基于优化的流水线中的奖励目标,VLM3D在标准基准上显著优于现有方法;(2) 作为前馈流水线的测试时引导模块,能主动修正SOTA原生3D模型迭代采样过程中的严重空间错误。VLM3D为将VLM丰富的、以语言为基础的语义与空间理解注入多样化的3D生成流水线提供了原则性且可泛化的路径。

原文摘要 · Abstract (English)

Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still face two fundamental limitations. First, they struggle with coarse semantic alignment, often failing to capture fine-grained prompt details. Second, they lack robust 3D spatial understanding, leading to geometric inconsistencies and catastrophic failures in part assembly and spatial relationships. To address these challenges, we propose VLM3D, a general framework that repurposes large vision-language models (VLMs) as powerful, differentiable semantic and spatial critics. Our core contribution is a dual-query critic signal derived from the VLM's Yes or No log-odds, which assesses both semantic fidelity and geometric coherence. We demonstrate the generality of this guidance signal across two distinct paradigms: (1) As a reward objective for optimization-based pipelines, VLM3D significantly outperforms existing methods on standard benchmarks. (2) As a test-time guidance module for feed-forward pipelines, it actively steers the iterative sampling process of SOTA native 3D models to correct severe spatial errors. VLM3D establishes a principled and generalizable path to inject the VLM's rich, language-grounded understanding of both semantics and space into diverse 3D generative pipelines.

3D生成视觉语言模型空间理解语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。