用结构能量引导采样,让文本生成3D模型多视角一致
Structural Energy-Guided Sampling for View-Consistent Text-to-3D
- 在采样阶段引入主成分分析的结构能量,引导几何走向正确视角
- 相比基线方法,显著减少视角偏差导致的几何坍缩与重复
- 无需训练或修改权重,可直接接入现有生成流程
文本生成3D模型常面临‘双面人’问题:正面看起来正确,但侧面出现重复或扭曲几何。我们归因于2D扩散先验中的视角偏差,并传播至3D优化过程。为此提出结构能量引导采样(SEGS),一种无需训练、即插即用的框架,完全在采样阶段实现多视角一致性。SEGS在中间U-Net特征的主成分子空间中定义结构能量,并将其梯度注入去噪轨迹,引导几何朝预期视角发展,同时保持外观保真度。该方法无缝集成至SDS/VSD流程,显著降低双面人伪影,提升几何对齐与视角一致性,且无需重新训练或修改权重。
原文摘要 · Abstract (English)
Text-to-3D generation often suffers from the Janus problem, where objects look correct from the front but collapse into duplicated or distorted geometry from other angles. We attribute this failure to viewpoint bias in 2D diffusion priors, which propagates into 3D optimization. To address this, we propose Structural Energy-Guided Sampling (SEGS), a training-free, plug-and-play framework that enforces multi-view consistency entirely at sampling time. SEGS defines a structural energy in a PCA subspace of intermediate U-Net features and injects its gradients into the denoising trajectory, steering geometry toward the intended viewpoint while preserving appearance fidelity. Integrated seamlessly into SDS/VSD pipelines, SEGS significantly reduces Janus artifacts, achieving improved geometric alignment and viewpoint consistency without retraining or weight modification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。