arXiv:2510.03874cs.CV2025-10被引 4

首个4D数字人质量评估数据集与多模态评分模型

DHQA-4D: Perceptual Quality Assessment of Dynamic 4D Digital Human

  • 构建包含1920个失真4D人像的评估数据集
  • 多模态融合特征提升对纹理/非纹理4D网格的评分精度
  • 适合数字人、虚拟现实、动画制作领域研究者使用

随着3D扫描与重建技术的快速发展,基于4D网格的动态数字人形象日益普及,可广泛应用于游戏制作、动画生成和远程沉浸式通信。然而,在采集、压缩和传输过程中,这些4D人像网格容易受到多种噪声干扰,影响用户体验。因此,动态4D数字人的质量评估变得尤为重要。本文首次提出大规模动态数字人质量评估数据集DHQA-4D,包含32个高精度真实扫描的4D人像序列、1920个经11种纹理失真处理的带纹理4D网格,以及对应的带纹理和无纹理平均意见得分(MOS)。基于该数据集,我们分析了不同失真类型对人眼感知的影响。此外,提出DynaMesh-Rater,一种基于大型多模态模型(LMM)的新方法,能够同时评估带纹理和无纹理4D网格。该方法从投影2D视频提取视觉特征,从裁剪视频片段提取运动特征,从4D网格提取几何特征,综合多维信息,并通过LoRA指令微调训练模型预测质量分数。在DHQA-4D上的大量实验表明,DynaMesh-Rater优于现有评估方法。

原文摘要 · Abstract (English)

With the rapid development of 3D scanning and reconstruction technologies, dynamic digital human avatars based on 4D meshes have become increasingly popular. A high-precision dynamic digital human avatar can be applied to various fields such as game production, animation generation, and remote immersive communication. However, these 4D human avatar meshes are prone to being degraded by various types of noise during the processes of collection, compression, and transmission, thereby affecting the viewing experience of users. In light of this fact, quality assessment of dynamic 4D digital humans becomes increasingly important. In this paper, we first propose a large-scale dynamic digital human quality assessment dataset, DHQA-4D, which contains 32 high-quality real-scanned 4D human mesh sequences, 1920 distorted textured 4D human meshes degraded by 11 textured distortions, as well as their corresponding textured and non-textured mean opinion scores (MOSs). Equipped with DHQA-4D dataset, we analyze the influence of different types of distortion on human perception for textured dynamic 4D meshes and non-textured dynamic 4D meshes. Additionally, we propose DynaMesh-Rater, a novel large multimodal model (LMM) based approach that is able to assess both textured 4D meshes and non-textured 4D meshes. Concretely, DynaMesh-Rater elaborately extracts multi-dimensional features, including visual features from a projected 2D video, motion features from cropped video clips, and geometry features from the 4D human mesh to provide comprehensive quality-related information. Then we utilize a LMM model to integrate the multi-dimensional features and conduct a LoRA-based instruction tuning technique to teach the LMM model to predict the quality scores. Extensive experimental results on the DHQA-4D dataset demonstrate the superiority of our DynaMesh-Rater method over previous quality assessment methods.

4D数字人质量评估多模态图像质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。