构建首个包含口腔与眼球的全身头像生成模型
GNM Head: A Generative aNthropometric Model of the human head

- 基于高精度3D扫描与人工建模数据,构建完整头面部参数化模型
- 在人脸扫描拟合任务中达到当前最佳性能(SotA)
- 适合影视动画、虚拟人生成与医学建模等需要高精度头部结构的研究
人体头部参数化模型是计算机视觉与图形学中用于动画、渲染和重建的重要工具。近年来,这类模型也成为生成式大视觉模型中的关键条件信号,实现对生成图像的空间精确控制。然而,现有公开模型通常局限于外部几何结构,忽略口腔与眼内结构,且因输入数据质量较低导致几何精度不足。本文提出新型参数化模型——生成型解剖模型(GNM),其名称为“基因组”的谐音。该模型涵盖头、脸、颈部、眼球、牙齿与舌头,并基于大规模高分辨率3D扫描数据集及高质量专业人工样本构建。报告详细说明了数据来源、模型架构(含眼部与口内结构专用子模型),并在目标3D人脸扫描拟合任务中展示其达到当前最优性能。为促进社区创新,完整GNM框架已公开发布。
原文摘要 · Abstract (English)
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery. However, existing publicly available models are typically limited in anatomical scope, modeling only outer geometry while ignoring intra-oral and ocular structures, and frequently suffer from reduced geometric quality stemming from low-fidelity input datasets. In this report we introduce a new parametric model dubbed Generative aNthropometric Model (GNM), named as a homophone of the human genome. GNM encompasses the head, face, neck, eyeballs, teeth, and tongue, and it is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy specific artist-made samples. This report details the data provenance, the model architecture including the specialized sub-models for the ocular and intra-oral structures, and shows its SotA performance on fitting target 3D face scans. To foster community innovation, the complete GNM framework is made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。