首个通用文本生成3D综合评测基准,解决数据与评估短板。
GT23D-Bench: A Comprehensive General Text-to-3D Generation Benchmark
- 构建40万件高质量3D资产数据集,含超7000万视觉样本和百万级层级描述。
- 提出10项多维度评测指标,对齐人类判断的相关性显著优于现有方法。
- 适用于评估通用文本生成3D模型,助力研究者发现共性缺陷与优化方向。
文本生成3D(T23D)已成为关键视觉生成任务,旨在从文本描述中合成3D内容。当前研究正从每场景优化的T23D转向仅需一个预训练模型即可生成多样内容的通用T23D(GT23D),以实现更广泛、高效的3D生成。然而,GT23D仍受限于两大核心挑战:高质量大规模训练数据缺失,以及评估指标忽视3D固有属性。现有数据集普遍存在标注不全、组织混乱、质量不一问题,而评估多依赖2D图像-文本相似度或评分,未能充分检验3D几何完整性与语义相关性。为此,我们推出首个专为GT23D设计的综合性基准GT23D-Bench。首先,构建包含40万3D资产、7000万+视觉样本及100万+层级描述的高质量数据集,促进鲁棒语义学习。其次,提出涵盖10项指标的综合评估体系,覆盖文本-3D对齐与3D视觉质量多个层面。关键的是,实验证明新指标与人类判断的相关性显著高于现有方法。基于该基准对8个领先GT23D模型的深度分析,揭示了当前模型能力与共性失效模式。GT23D-Bench将公开发布,推动严谨可复现的研究。
原文摘要 · Abstract (English)
Text-to-3D (T23D) generation has emerged as a crucial visual generation task, aiming at synthesizing 3D content from textual descriptions. Studies of this task are currently shifting from per-scene T23D, which requires optimization of the model for every content generated, to General T23D (GT23D), which requires only one pre-trained model to generate different content without re-optimization, for more generalized and efficient 3D generation. Despite notable advancements, GT23D is severely bottlenecked by two interconnected challenges: the lack of high-quality, large-scale training data and the prevalence of evaluation metrics that overlook intrinsic 3D properties. Existing datasets often suffer from incomplete annotations, noisy organization, and inconsistent quality, while current evaluations rely heavily on 2D image-text similarity or scoring, failing to thoroughly assess 3D geometric integrity and semantic relevance. To address these fundamental gaps, we introduce GT23D-Bench, the first comprehensive benchmark specifically designed for GT23D training and evaluation. We first construct a high-quality dataset of 400K 3D assets, featuring diverse visual annotations (70M+ visual samples) and multi-granularity hierarchical captions (1M+ descriptions) to foster robust semantic learning. Second, we propose a comprehensive evaluation suite with 10 metrics assessing both text-3D alignment and 3D visual quality at multiple levels. Crucially, we demonstrate through rigorous experiments that our proposed metrics exhibit significantly higher correlation with human judgment compared to existing methods. Our in-depth analysis of eight leading GT23D models using this benchmark provides the community with critical insights into current model capabilities and their shared failure modes. GT23D-Bench will be publicly available to facilitate rigorous and reproducible research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。