构建细粒度3D生成评估基准,提出两阶段评分模型提升预测准确性。
Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- 设计包含12子项的组合式提示框架,生成3600个纹理网格数据集
- 通过12.96万次人工评分建立高质量主观评价基准
- 两阶段训练模型更贴近人类判断,可作为生成模型优化奖励函数
Text-to-3D(T23D)生成模型近年来实现了从文本提示生成多样且高保真3D资产的能力。然而,现有挑战限制了可靠T23D质量评估(T23DQA)的发展:首先,现有基准过时、分散且粗粒度,难以支持细粒度度量训练;其次,当前客观度量存在固有设计缺陷,导致特征提取不具代表性且鲁棒性下降。为此,我们提出T23D-CompBench,一个用于组合式T23D生成的综合性基准。定义五个组件与十二个子组件构成组合提示,由十种先进生成模型生成3,600个带纹理网格。通过大规模主观实验收集129,600次跨视角人类评分。基于该基准,进一步提出Rank2Score,一种两阶段训练的T23DQA评估器。第一阶段通过监督对比回归与课程学习增强成对训练;第二阶段利用平均意见分精炼预测,使其更贴近人类判断。大量实验与下游应用表明,Rank2Score在多维度上持续优于现有度量,并可作为生成模型优化的奖励函数。项目地址:https://cbysjtu.github.io/Rank2Score/
原文摘要 · Abstract (English)
Recent advances in Text-to-3D (T23D) generative models have enabled the synthesis of diverse, high-fidelity 3D assets from textual prompts. However, existing challenges restrict the development of reliable T23D quality assessment (T23DQA). First, existing benchmarks are outdated, fragmented, and coarse-grained, making fine-grained metric training infeasible. Moreover, current objective metrics exhibit inherent design limitations, resulting in non-representative feature extraction and diminished metric robustness. To address these limitations, we introduce T23D-CompBench, a comprehensive benchmark for compositional T23D generation. We define five components with twelve sub-components for compositional prompts, which are used to generate 3,600 textured meshes from ten state-of-the-art generative models. A large-scale subjective experiment is conducted to collect 129,600 reliable human ratings across different perspectives. Based on T23D-CompBench, we further propose Rank2Score, an effective evaluator with two-stage training for T23DQA. Rank2Score enhances pairwise training via supervised contrastive regression and curriculum learning in the first stage, and subsequently refines predictions using mean opinion scores to achieve closer alignment with human judgments in the second stage. Extensive experiments and downstream applications demonstrate that Rank2Score consistently outperforms existing metrics across multiple dimensions and can additionally serve as a reward function to optimize generative models. The project is available at https://cbysjtu.github.io/Rank2Score/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。