arXiv:2410.07658cs.CV2024-10被引 3

提出新框架,同时提升文本到3D生成的语义一致性和多视角一致性。

SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation

  • 引入正交注意力机制学习三平面先验,确保多视角几何一致。
  • 通过注意力对齐实现文本语义与三平面表征的跨视角一致合成。
  • 在多视角一致性上达到新SOTA,兼顾语义保真度,适合3D内容生成研究者。

通用文本到3D(GT23D)生成对构建多样化物体与场景的3D内容至关重要,但面临两大挑战:一是输入文本与生成3D模型间的语义一致性,二是不同视角间的一致性。现有方法通常仅解决其中一问题,导致语义保真度和结构连贯性不足。为此,我们提出SeMv-3D框架,联合优化语义对齐与多视角一致性。核心是三平面先验学习(TPL),通过专用正交注意力机制捕捉三个正交平面间的空间对应关系,确保视角间几何一致性;同时提出基于先验的三平面语义对齐(SAT),利用注意力特征对齐强化文本语义与三平面表示的对应关系,实现一致的任意视角合成。大量实验表明,该方法在多视角一致性上达到新SOTA,同时在语义一致性上保持竞争力,显著平衡并超越了两个维度的表现,树立了领域新基准。

原文摘要 · Abstract (English)

General Text-to-3D (GT23D) generation is crucial for creating diverse 3D content across objects and scenes, yet it faces two key challenges: 1) ensuring semantic consistency between input text and generated 3D models, and 2) maintaining multi-view consistency across different perspectives within 3D. Existing approaches typically address only one of these challenges, often leading to suboptimal results in semantic fidelity and structural coherence. To overcome these limitations, we propose SeMv-3D, a novel framework that jointly enhances semantic alignment and multi-view consistency in GT23D generation. At its core, we introduce Triplane Prior Learning (TPL), which effectively learns triplane priors by capturing spatial correspondences across three orthogonal planes using a dedicated Orthogonal Attention mechanism, thereby ensuring geometric consistency across viewpoints. Additionally, we present Prior-based Semantic Aligning in Triplanes (SAT), which enables consistent any-view synthesis by leveraging attention-based feature alignment to reinforce the correspondence between textual semantics and triplane representations. Extensive experiments demonstrate that our method sets a new state-of-the-art in multi-view consistency, while maintaining competitive performance in semantic consistency compared to methods focused solely on semantic alignment. These results emphasize the remarkable ability of our approach to effectively balance and excel in both dimensions, establishing a new benchmark in the field.

文本生成3D多视角一致三平面表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。