用视觉语言模型知识控制3D生成背面,让想象更可控
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
- 通过注入视觉语言模型的隐状态,将文本指令转化为3D生成指导
- 使未观测区域的生成从随机变为语义可控,提升结构合理性
- 适合需要精准控制3D生成结果的研究者与设计师
近期3D生成技术在保真度和几何细节上取得进展,但单视图观测的固有歧义及有限3D训练数据导致缺乏稳健的全局结构先验,使得现有模型生成的未见区域常呈随机性且难以控制,可能偏离用户意图或产生不合理几何结构。本文提出Know3D框架,通过潜在隐状态注入,将多模态大语言模型中的丰富知识融入3D生成过程,实现对3D资产背面的语义可控生成。采用基于视觉语言模型(VLM)与扩散模型的联合架构:VLM负责语义理解与引导,扩散模型作为桥梁将语义知识传递至3D生成模型。该方法成功弥合抽象文本指令与未观测区域几何重建之间的差距,将传统的随机背面幻觉转变为语义可控过程,为未来3D生成模型提供新方向。
原文摘要 · Abstract (English)
Recent advances in 3D generation have improved the fidelity and geometric details of synthesized 3D assets. However, due to the inherent ambiguity of single-view observations and the lack of robust global structural priors caused by limited 3D training data, the unseen regions generated by existing models are often stochastic and difficult to control, which may sometimes fail to align with user intentions or produce implausible geometries. In this paper, we propose Know3D, a novel framework that incorporates rich knowledge from multimodal large language models into 3D generative processes via latent hidden-state injection, enabling language-controllable generation of the back-view for 3D assets. We utilize a VLM-diffusion-based model, where the VLM is responsible for semantic understanding and guidance. The diffusion model acts as a bridge that transfers semantic knowledge from the VLM to the 3D generation model. In this way, we successfully bridge the gap between abstract textual instructions and the geometric reconstruction of unobserved regions, transforming the traditionally stochastic back-view hallucination into a semantically controllable process, demonstrating a promising direction for future 3D generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。