解决文本生成3D模型时的结构错乱和细节断裂问题。
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation

- 用语义视角一致性约束和流形特征连续性模块协同优化
- 在多个视图下提升3D物体整体结构与微细节的连贯性
- 适合关注高质量3D生成、特别是几何一致性的研究者
随着元宇宙、虚拟现实和数字孪生的发展,文本生成3D成为学术与产业热点。当前主流方法基于利用2D扩散先验的得分蒸馏采样(SDS),但受2D先验视角偏差及高无分类指导(CFG)带来的梯度噪声影响,仍存在宏观拓扑不一致(如双面人问题)和微观几何断续的问题。为此,我们提出MOC-3D,基于几何流形与语义视角一致性设计。在ScaleDreamer框架基础上,引入语义视角顺序约束模块与流形特征连续性模块。前者利用CLIP先验,在不同视角间对语义得分表示施加单调性排序约束,有效引导3D物体全局拓扑结构;后者在对称正定(SPD)流形上使用黎曼度量,通过统计分布距离衡量特征变化,促进多视角下微纹理的平滑演化。两者协同实现宏-微观同步优化,显著提升生成3D对象的结构一致性与细节连续性。
原文摘要 · Abstract (English)
With the burgeoning development of fields such as the Metaverse, Virtual Reality (VR), and Digital Twins, text-to-3D generation has emerged as a research hotspot in both academia and industry. Currently, optimization methods based on Score Distillation Sampling (SDS) utilizing 2D diffusion priors have become the mainstream technological paradigm in this field. However, due to the view bias of 2D priors and the mode-seeking ambiguity combined with gradient noise induced by high Classifier-Free Guidance (CFG), these methods still suffer from macro-topological inconsistency (e.g., the Janus problem) and micro-geometric discontinuity. To address these challenges, we propose MOC-3D, a text-to-3D generation method based on geometric manifold and semantic view-order consistency. Built upon the ScaleDreamer framework, our method incorporates a Semantic View-Order Constraint Module and a Manifold-based Feature Continuity Module. The former aims to rectify macro-topological inconsistency, while the latter focuses on eliminating micro-geometric discontinuity. Specifically, the Semantic View-Order Constraint Module leverages the prior knowledge of CLIP to impose a Monotonicity Rank Constraint on semantic score representations across different views, thereby providing effective guidance for the global topological structure of 3D objects. Meanwhile, the Manifold-based Feature Continuity Module employs the Riemannian Metric on the Symmetric Positive Definite (SPD) manifold. By measuring the distance of feature statistical distributions in the Riemannian space, it promotes the smooth evolution and continuity of micro-textures across multi-views in a statistical sense. Under the macro-micro synergistic optimization of these two modules, our model can simultaneously improve macro-structural consistency and micro-detail continuity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。