arXiv:2602.13818cs.CVcs.LG2026-02被引 2

用视图感知的3D编码器提升文本生成3D模型的质量与对齐度。

VAR-3D: View-aware Auto-Regressive Model for Text-to-3D Generation via a 3D Tokenizer

  • 设计视图感知的3D VQ-VAE,将3D几何转为离散标记。
  • 引入渲染监督训练,使生成结果更贴合输入文本。
  • 适合需要高保真3D生成与文本对齐的研究者。

自回归变换器在生成建模中取得显著进展,但文本到3D生成仍面临挑战,主要源于学习离散3D表示的瓶颈。现有方法在编码阶段常出现信息丢失,导致量化前的表征失真,且向量量化进一步放大该问题,最终降低文本条件3D形状的几何一致性。此外,传统的两阶段训练范式在重建与文本条件自回归生成间存在目标不匹配。为此,我们提出视图感知自回归3D(VAR-3D),结合视图感知3D向量量化变分自编码器(VQ-VAE),将复杂3D结构转换为离散标记。同时,引入渲染监督训练策略,将离散标记预测与视觉重建耦合,促使生成过程更好保持视觉保真度与结构一致性。实验表明,VAR-3D在生成质量与文本-3D对齐方面显著优于现有方法。

原文摘要 · Abstract (English)

Recent advances in auto-regressive transformers have achieved remarkable success in generative modeling. However, text-to-3D generation remains challenging, primarily due to bottlenecks in learning discrete 3D representations. Specifically, existing approaches often suffer from information loss during encoding, causing representational distortion before the quantization process. This effect is further amplified by vector quantization, ultimately degrading the geometric coherence of text-conditioned 3D shapes. Moreover, the conventional two-stage training paradigm induces an objective mismatch between reconstruction and text-conditioned auto-regressive generation. To address these issues, we propose View-aware Auto-Regressive 3D (VAR-3D), which intergrates a view-aware 3D Vector Quantized-Variational AutoEncoder (VQ-VAE) to convert the complex geometric structure of 3D models into discrete tokens. Additionally, we introduce a rendering-supervised training strategy that couples discrete token prediction with visual reconstruction, encouraging the generative process to better preserve visual fidelity and structural consistency relative to the input text. Experiments demonstrate that VAR-3D significantly outperforms existing methods in both generation quality and text-3D alignment.

3D生成自回归文本生成向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。