arXiv:2605.19876cs.CV2026-05

解决文本生成3D的视角不一致问题,无需训练即可提升多视角一致性。

Structural Energy Guidance for View-Consistent Text-to-3D Generation

论文配图:Structural Energy Guidance for View-Consistent Text-to-3D Generation
图 1 · 摘自论文原文
  • 在U-Net特征主成分空间构建结构能量,通过梯度引导去噪过程。
  • 平均降低10%的Janus率,多个基线模型的视图一致性评分提升。
  • 无需重训练,可直接接入现有生成流程,适合高质量3D内容生成场景。

基于扩散模型的文本生成3D常面临视角不一致的‘双面人’问题。本文识别出2D扩散先验中的视角偏差是主要原因,提出无需训练、即插即用的结构能量引导采样(SEGS)框架。SEGS在U-Net特征的主成分子空间构建结构能量,并将其梯度注入去噪过程,可无缝集成至SDS/VSD流程中。实验表明,该方法平均降低约10%的Janus率,在DreamFusion、Magic3D和LucidDreamer等多个基线上均提升视图一致性(View-CS)得分。该方法有效缓解视角伪影,同时保持外观保真度,为高质量文本生成3D提供灵活解决方案。

原文摘要 · Abstract (English)

Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D diffusion priors as the main cause and proposes Structural Energy-Guided Sampling (SEGS), a training-free and plug-and-play framework to improve multi-view consistency. SEGS constructs a structural energy in the PCA subspace of U-Net features and injects its gradient into the denoising process. It can be easily integrated into SDS/VSD pipelines without retraining. Experiments show that SEGS reduces the Janus Rate by about 10% on average and improves View-CS scores across multiple baselines, including DreamFusion, Magic3D, and LucidDreamer. This method effectively alleviates viewpoint artifacts while preserving appearance fidelity, providing a flexible solution for high-quality text-to-3D content generation.

3D生成扩散模型一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。