arXiv:2502.08977cs.CV2025-02被引 2

用对比偏好优化提升文本生成3D人体的精准度

Text-driven 3D Human Generation via Contrastive Preference Optimization

  • 引入正负提示的对比偏好,增强文本与3D模型对齐
  • 在长文本输入下显著提升纹理真实感与视觉一致性
  • 适合需要高精度3D人体生成的研究者和开发者

基于分数蒸馏采样(SDS)的3D人体生成近期取得进展,但面对长而复杂的文本描述仍存在对齐难题。为此,本文提出一种新框架,引入人类偏好模型指导的对比偏好机制,结合正负提示优化对齐效果。设计偏好优化模块融合多模型,全面捕捉文本特征;引入否定偏好模块,通过静态-动态否定提示缓解无关细节过度优化问题,有效防止“奖励劫持”。大量实验表明,该方法达到当前最优性能,在长复杂文本输入下显著提升纹理真实感与视觉对齐度。

原文摘要 · Abstract (English)

Recent advances in Score Distillation Sampling (SDS) have improved 3D human generation from textual descriptions. However, existing methods still face challenges in accurately aligning 3D models with long and complex textual inputs. To address this challenge, we propose a novel framework that introduces contrastive preferences, where human-level preference models, guided by both positive and negative prompts, assist SDS for improved alignment. Specifically, we design a preference optimization module that integrates multiple models to comprehensively capture the full range of textual features. Furthermore, we introduce a negation preference module to mitigate over-optimization of irrelevant details by leveraging static-dynamic negation prompts, effectively preventing ``reward hacking". Extensive experiments demonstrate that our method achieves state-of-the-art results, significantly enhancing texture realism and visual alignment with textual descriptions, particularly for long and complex inputs.

3D生成文本驱动偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。