用2D奖励模型让3D生成更符合人类偏好,无需3D训练数据。
Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
- 通过类无分类器引导机制,将2D奖励信号转化为3D生成的对齐信号。
- 在Aesthetic、CLIP Score等指标上优于现有方法,且可无缝接入不同生成流程。
- 适合追求高质量3D生成与人类偏好对齐的研究者和开发者。
人类偏好对齐是扩散模型在文本到3D生成中一个关键但研究不足的挑战。现有方法通常需要特定任务的微调,在数据稀缺的3D领域面临巨大障碍。为此,我们提出偏好得分蒸馏(Preference Score Distillation, PSD),一种基于优化的框架,利用预训练的2D奖励模型实现无需3D训练数据的人类对齐文本到3D合成。核心洞察在于:由于奖励模型训练中缺乏噪声样本,直接应用2D奖励梯度会破坏去噪过程。借鉴条件扩散模型中朴素分类器引导的问题,我们将偏好对齐重新构想为类无分类器引导(CFG)机制,并通过隐式奖励模型实现。此外,考虑到冻结的预训练扩散模型限制性能,我们引入自适应策略,联合优化偏好得分与负向文本嵌入。在优化过程中引入CFG,使负向文本嵌入在线更新,动态提升对齐效果。据我们所知,这是首个在得分蒸馏框架下将人类偏好对齐与CFG理论结合的工作。实验表明,PSD在美学指标、多种流程的无缝集成以及强可扩展性方面均表现出优越性。
原文摘要 · Abstract (English)
Human preference alignment presents a critical yet underexplored challenge for diffusion models in text-to-3D generation. Existing solutions typically require task-specific fine-tuning, posing significant hurdles in data-scarce 3D domains. To address this, we propose Preference Score Distillation (PSD), an optimization-based framework that leverages pretrained 2D reward models for human-aligned text-to-3D synthesis without 3D training data. Our key insight stems from the incompatibility of pixel-level gradients: due to the absence of noisy samples during reward model training, direct application of 2D reward gradients disturbs the denoising process. Noticing that similar issue occurs in the naive classifier guidance in conditioned diffusion models, we fundamentally rethink preference alignment as a classifier-free guidance (CFG)-style mechanism through our implicit reward model. Furthermore, recognizing that frozen pretrained diffusion models constrain performance, we introduce an adaptive strategy to co-optimize preference scores and negative text embeddings. By incorporating CFG during optimization, online refinement of negative text embeddings dynamically enhances alignment. To our knowledge, we are the first to bridge human preference alignment with CFG theory under score distillation framework. Experiments demonstrate the superiority of PSD in aesthetic metrics, seamless integration with diverse pipelines, and strong extensibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。