构建首个大规模3D生成质量评估数据集与模型
Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and Model
- 建立969个3D资产的多维评价数据集
- 实现质量、真实感、文本对应度三维度评估
- 适配AI生成内容质量评测的科研与工业场景
近期文本到图像(T2I)生成技术的发展推动了文本到3D资产(T23DA)生成的兴起,该技术利用预训练的2D文本到图像扩散模型实现文本到3D资产的合成。尽管T23DA生成日益流行,其评估体系尚未得到充分研究。鉴于生成资产间存在显著质量差异,亟需与人类主观判断对齐的质量评估模型。为此,本文从主观与客观双视角展开全面研究,首次构建迄今最大的文本到3D资产质量评估数据库——AIGC-T23DAQA,包含由170个提示词通过6种主流模型生成的969个经验证的3D资产,并提供质量、真实性及文本-资产对应度三方面的主观评分。基于该数据库,我们建立了综合性基准,并设计了一种有效T23DAQA模型,可分别从上述三个维度评估生成的3D资产。
原文摘要 · Abstract (English)
Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。