首个统一3D理解与生成的框架,支持图像转3D场景和空间问答。
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- 用语言模型解析文本与3D表示,结合潜在扩散模型生成高质量3D内容。
- 在3D生成与空间问答任务上优于现有方法,支持任意视角变换生成。
- 提出几何-语义联合学习策略,提升视觉编码器的空间感知能力。
尽管近期统一架构在图像理解与生成上取得显著进展,但3D任务的融合仍具挑战且研究不足。本文提出UniUGG,首个面向3D模态的统一理解与生成框架。该框架利用大语言模型(LLM)理解并解码句子与3D表示。核心创新是基于潜在扩散模型的时空解码器,可依据参考图像与任意视角变换生成高质量3D场景,并支持空间视觉问答(VQA)任务。此外,我们设计了几何-语义学习策略预训练视觉编码器,联合捕捉输入的语义与几何线索,增强空间理解与生成能力。大量实验表明,该方法在视觉表征、空间理解与3D生成方面均具优势。
原文摘要 · Abstract (English)
Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexplored. In this paper, we introduce UniUGG, the first unified understanding and generation framework for 3D modalities. Our unified framework employs an LLM to comprehend and decode sentences and 3D representations. At its core, we propose a spatial decoder leveraging a latent diffusion model to generate high-quality 3D representations. This allows for the generation and imagination of 3D scenes based on a reference image and an arbitrary view transformation, while remaining supports for spatial visual question answering (VQA) tasks. Additionally, we propose a geometric-semantic learning strategy to pretrain the vision encoder. This design jointly captures the input's semantic and geometric cues, enhancing both spatial understanding and generation. Extensive experimental results demonstrate the superiority of our method in visual representation, spatial understanding, and 3D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。