解决文本生成3D时多视角不一致问题,提升生成质量。
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
- 通过视图解耦模块消除提示中的视角偏差
- 利用相似性损失保证无条件项的几何一致性
- 适用于多种3D表示和扩散框架,适合3D生成研究者
零样本文本到3D生成技术已推动3D内容创作革新,实现从文本直接生成3D模型。现有方法虽借助3D高斯溅射与得分蒸馏,结合预训练文本到图像(T2I)模型增强多视角渲染,但受限于T2I先验的固有视角偏见,导致3D生成不一致,尤其表现为多面矛盾的‘雅努斯问题’。为此,我们提出ConsDreamer,通过优化得分蒸馏过程中的条件与无条件项来缓解视角偏差:(1) 视图解耦模块(VDM)分离无关视图成分,注入精确视图控制,消除条件提示中的视角偏见;(2) 基于相似性的部分顺序损失,使无条件项的余弦相似度与方位角关系对齐,强化几何一致性。大量实验表明,ConsDreamer可无缝集成至多种3D表示与得分蒸馏范式,有效缓解多面雅努斯问题。
原文摘要 · Abstract (English)
Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score distillation to enhance multi-view rendering through pre-trained text-to-image (T2I) models, they suffer from inherent prior view biases in T2I priors. These biases lead to inconsistent 3D generation, particularly manifesting as the multi-face Janus problem, where objects exhibit conflicting features across views. To address this fundamental challenge, we propose ConsDreamer, a novel method that mitigates view bias by refining both the conditional and unconditional terms in the score distillation process: (1) a View Disentanglement Module (VDM) that eliminates viewpoint biases in conditional prompts by decoupling irrelevant view components and injecting precise view control; and (2) a similarity-based partial order loss that enforces geometric consistency in the unconditional term by aligning cosine similarities with azimuth relationships. Extensive experiments demonstrate that ConsDreamer can be seamlessly integrated into various 3D representations and score distillation paradigms, effectively mitigating the multi-face Janus problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。