arXiv:2506.15684cs.GRcs.CV2025-06NeurIPS被引 9

用2D奖励高效优化3D生成模型,让机器造物更贴近人类审美。

Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

  • 通过2D图像奖励引导3D扩散模型,无需3D标注
  • 仅用少量迭代就大幅提升奖励得分并减少遗忘
  • 适合需要快速对齐人类偏好的3D内容生成场景

高质量且逼真的3D资产生成仍是3D视觉与计算机图形学中的长期挑战。尽管当前先进的生成模型(如扩散模型)在3D生成方面取得显著进展,但仍难以达到人工设计水准,主要受限于指令遵循能力不足、人类偏好对齐困难,以及真实纹理、几何结构和物理属性生成不佳。本文提出Nabla-R2D3,一种基于2D奖励信号的高效3D原生扩散模型强化学习对齐框架。该方法建立在最近提出的Nabla-GFlowNet基础上,以严谨方式将评分函数与奖励梯度匹配,实现奖励微调。实验表明,相比传统微调基线(易发散或出现奖励欺骗),Nabla-R2D3在少数微调步骤内即可稳定获得更高奖励,并有效缓解先验遗忘问题。

原文摘要 · Abstract (English)

Generating high-quality and photorealistic 3D assets remains a longstanding challenge in 3D vision and computer graphics. Although state-of-the-art generative models, such as diffusion models, have made significant progress in 3D generation, they often fall short of human-designed content due to limited ability to follow instructions, align with human preferences, or produce realistic textures, geometries, and physical attributes. In this paper, we introduce Nabla-R2D3, a highly effective and sample-efficient reinforcement learning alignment framework for 3D-native diffusion models using 2D rewards. Built upon the recently proposed Nabla-GFlowNet method, which matches the score function to reward gradients in a principled manner for reward finetuning, our Nabla-R2D3 enables effective adaptation of 3D diffusion models using only 2D reward signals. Extensive experiments show that, unlike vanilla finetuning baselines which either struggle to converge or suffer from reward hacking, Nabla-R2D3 consistently achieves higher rewards and reduced prior forgetting within a few finetuning steps.

3D生成扩散模型强化学习奖励对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。