arXiv:2511.08203cs.CV2025-11中稿 · EurIPS 2025 Worksh…

发现3D生成模型对视角有强依赖,旋转输入会显著降低效果。

Twist and Compute: The Cost of Pose in 3D Generative Diffusion

  • 用2D旋转测试揭示主流3D生成模型存在视角偏见
  • 旋转后性能下降,但加轻量CNN可恢复效果
  • 提示应设计更对称的模块化架构,而非单纯堆规模

尽管大规模图像到3D生成模型取得了显著成果,但其归纳偏置仍不透明。我们识别出图像条件3D生成模型的一个重要局限:强烈的规范视图偏见。通过使用简单2D旋转的受控实验,我们发现当前最先进的Hunyuan3D 2.0模型在输入旋转时难以泛化,性能明显下降。我们证明,通过添加一个轻量级CNN来检测并校正输入方向,可恢复模型性能,且无需修改生成主干。这一发现提出一个重要开放问题:规模是否足够?还是应追求模块化、对称感知的设计?

原文摘要 · Abstract (English)

Despite their impressive results, large-scale image-to-3D generative models remain opaque in their inductive biases. We identify a significant limitation in image-conditioned 3D generative models: a strong canonical view bias. Through controlled experiments using simple 2D rotations, we show that the state-of-the-art Hunyuan3D 2.0 model can struggle to generalize across viewpoints, with performance degrading under rotated inputs. We show that this failure can be mitigated by a lightweight CNN that detects and corrects input orientation, restoring model performance without modifying the generative backbone. Our findings raise an important open question: Is scale enough, or should we pursue modular, symmetry-aware designs?

3D生成扩散模型视角不变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。