用大模型评估3D形状真实感,无需参考标准。
SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation
- 将3D网格转为文本令牌,通过大模型判断真实感。
- 在16种算法生成的上千个形状上验证,与人评一致。
- 适合无真值时的3D内容质量评估,如游戏建模。
3D生成与重建技术广泛应用于游戏、影视等内容创作领域。随着应用扩展,对视觉真实感更高的3D形状需求日益增长。传统评估依赖真值对比,但在多数实际场景中,真实感并不依赖真值存在。本文提出一种无需参考的3D形状真实感度量方法——形状真实感对齐度量(SRAM),利用大语言模型(LLM)作为连接3D网格信息与真实感评价的桥梁。通过网格编码将3D形状映射至语言空间,并设计专用真实感解码器,使LLM输出对齐人类对真实感的感知。同时构建新数据集RealismGrading,包含16种不同算法在十余类物体上生成的形状,提供无真值的人工标注真实感评分。通过跨对象的k折交叉验证,实验表明该度量与人类感知高度相关,优于现有方法,具备良好泛化能力。
原文摘要 · Abstract (English)
3D generation and reconstruction techniques have been widely used in computer games, film, and other content creation areas. As the application grows, there is a growing demand for 3D shapes that look truly realistic. Traditional evaluation methods rely on a ground truth to measure mesh fidelity. However, in many practical cases, a shape's realism does not depend on having a ground truth reference. In this work, we propose a Shape-Realism Alignment Metric that leverages a large language model (LLM) as a bridge between mesh shape information and realism evaluation. To achieve this, we adopt a mesh encoding approach that converts 3D shapes into the language token space. A dedicated realism decoder is designed to align the language model's output with human perception of realism. Additionally, we introduce a new dataset, RealismGrading, which provides human-annotated realism scores without the need for ground truth shapes. Our dataset includes shapes generated by 16 different algorithms on over a dozen objects, making it more representative of practical 3D shape distributions. We validate our metric's performance and generalizability through k-fold cross-validation across different objects. Experimental results show that our metric correlates well with human perceptions and outperforms existing methods, and has good generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。