用统一模型从任意图像序列重建高保真4D人脸,提升追踪精度与稳定性。
Face Anything: 4D Face Reconstruction from Any Image Sequence
- 基于规范面部坐标预测,将动态重建转为统一前向建模问题。
- 对应点误差降低3倍,深度估计准确率提升16%,推理更快。
- 适合需要高精度人脸重建与追踪的虚拟人、影视特效应用。
从图像序列中精确重建和追踪动态人脸极具挑战,因非刚性变形、表情变化和视角差异同时发生,导致几何与对应关系估计存在显著歧义。本文提出一种基于规范面部点预测的统一方法,该表示将每个像素映射到共享规范空间中的归一化面部坐标。此公式将密集追踪与动态重建转化为规范重建问题,实现单次前馈模型下的时序一致几何与可靠对应关系。通过联合预测深度与规范坐标,本方法在单一架构内实现了精确深度估计、时序稳定重建、稠密3D几何及鲁棒面部点追踪。采用基于Transformer的模型,利用非刚性对齐至规范空间的多视角几何数据进行训练。在图像与视频基准测试中,实验表明本方法在重建与追踪任务上均达到当前最优性能,对应点误差降低约3倍,推理速度更快,深度准确率提升16%。结果验证了规范面部点预测作为统一前馈4D人脸重建有效基础的潜力。
原文摘要 · Abstract (English)
Accurate reconstruction and tracking of dynamic human faces from image sequences is challenging because non-rigid deformations, expression changes, and viewpoint variations occur simultaneously, creating significant ambiguity in geometry and correspondence estimation. We present a unified method for high-fidelity 4D facial reconstruction based on canonical facial point prediction, a representation that assigns each pixel a normalized facial coordinate in a shared canonical space. This formulation transforms dense tracking and dynamic reconstruction into a canonical reconstruction problem, enabling temporally consistent geometry and reliable correspondences within a single feed-forward model. By jointly predicting depth and canonical coordinates, our method enables accurate depth estimation, temporally stable reconstruction, dense 3D geometry, and robust facial point tracking within a single architecture. We implement this formulation using a transformer-based model that jointly predicts depth and canonical facial coordinates, trained using multi-view geometry data that non-rigidly warps into the canonical space. Extensive experiments on image and video benchmarks demonstrate state-of-the-art performance across reconstruction and tracking tasks, achieving approximately 3$\times$ lower correspondence error and faster inference than prior dynamic reconstruction methods, while improving depth accuracy by 16%. These results highlight canonical facial point prediction as an effective foundation for unified feed-forward 4D facial reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。