arXiv:2412.06614cs.CV2024-12AAAI被引 11

构建人类偏好对齐的3D生成评估框架,提升多视角扩散模型评价公平性。

MVReward: Better Aligning and Evaluating Multi-View Diffusion Models with Human Preferences

  • 基于16k张专家对比数据训练奖励模型MVReward,精准捕捉人类偏好
  • 提出MVP调优策略,显著提升多视角生成与人类审美的对齐度
  • 适用于评估和优化文本/图像驱动的3D生成模型,尤其适合研究者

近年来3D内容生成取得显著进展,但相应评估方法难以跟上。自动评估难以对齐人类偏好,且文本与图像驱动方法混合比较常导致不公平结果。本文提出一个全面框架,更好对齐并评估多视角扩散模型与人类偏好。首先从DALL·E和Objaverse收集并筛选标准化图像提示集,用于生成多视角资产,并通过系统化排序流程获得包含16,000条专家成对比较的人类标注数据集,进而训练出名为MVReward的奖励模型,以有效编码人类偏好。利用MVReward,可更公平透明地评估图像驱动3D方法。在此基础上,进一步提出即插即用的多视角偏好学习(MVP)策略。大量实验表明,MVReward可作为可靠度量,MVP能持续提升多视角扩散模型与人类偏好的对齐程度。

原文摘要 · Abstract (English)

Recent years have witnessed remarkable progress in 3D content generation. However, corresponding evaluation methods struggle to keep pace. Automatic approaches have proven challenging to align with human preferences, and the mixed comparison of text- and image-driven methods often leads to unfair evaluations. In this paper, we present a comprehensive framework to better align and evaluate multi-view diffusion models with human preferences. To begin with, we first collect and filter a standardized image prompt set from DALL$\cdot$E and Objaverse, which we then use to generate multi-view assets with several multi-view diffusion models. Through a systematic ranking pipeline on these assets, we obtain a human annotation dataset with 16k expert pairwise comparisons and train a reward model, coined MVReward, to effectively encode human preferences. With MVReward, image-driven 3D methods can be evaluated against each other in a more fair and transparent manner. Building on this, we further propose Multi-View Preference Learning (MVP), a plug-and-play multi-view diffusion tuning strategy. Extensive experiments demonstrate that MVReward can serve as a reliable metric and MVP consistently enhances the alignment of multi-view diffusion models with human preferences.

3D生成扩散模型偏好对齐评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。