让3D生成模型输出统一朝向,提升下游任务可用性
Orientation Matters: Making 3D Generative Models Orientation-Aligned
- 构建14,832个统一朝向的3D模型数据集Objaverse-OA
- 在1,008类物体上实现跨类别方向对齐生成
- 支持零样本朝向估计与箭头式旋转操作
人类能从单张图像直观感知物体形状与朝向,依赖于对标准姿态的强烈先验。然而现有3D生成模型因训练数据不一致,常产生朝向错乱的结果,限制了下游应用。为此,我们提出方向对齐的3D物体生成任务:从单图生成具有跨类别一致朝向的3D物体。为支持此任务,我们构建了包含14,832个3D模型、覆盖1,008个类别的Objaverse-OA数据集。基于该数据集,我们微调了两种代表性3D生成模型(基于多视角扩散与3D变分自编码器框架),使生成结果在未见物体和多种类别间具有良好泛化能力。实验表明,本方法优于事后对齐方案。此外,我们展示了下游应用:通过分析-合成实现零样本物体朝向估计,以及基于箭头的高效物体旋转操控。
原文摘要 · Abstract (English)
Humans intuitively perceive object shape and orientation from a single image, guided by strong priors about canonical poses. However, existing 3D generative models often produce misaligned results due to inconsistent training data, limiting their usability in downstream tasks. To address this gap, we introduce the task of orientation-aligned 3D object generation: producing 3D objects from single images with consistent orientations across categories. To facilitate this, we construct Objaverse-OA, a dataset of 14,832 orientation-aligned 3D models spanning 1,008 categories. Leveraging Objaverse-OA, we fine-tune two representative 3D generative models based on multi-view diffusion and 3D variational autoencoder frameworks to produce aligned objects that generalize well to unseen objects across various categories. Experimental results demonstrate the superiority of our method over post-hoc alignment approaches. Furthermore, we showcase downstream applications enabled by our aligned object generation, including zero-shot object orientation estimation via analysis-by-synthesis and efficient arrow-based object rotation manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。