arXiv:2412.08684cs.CVeess.IV2024-12CVPR被引 3

融合静态模型与动态帧,实现真实感3D人脸视频的稳定重建。

Coherent3D: Coherent 3D Portrait Video Reconstruction via Triplane Fusion

  • 用参考图构建3D先验,结合每帧输入动态融合外观特征。
  • 在真实和野外场景下均达到顶尖3D重建与时间一致性表现。
  • 适合追求高保真3D虚拟人直播与远程参会的开发者使用。

单图像3D人脸重建的最新进展使单摄像头实时流式传输3D人脸视频成为可能,推动了远程存在系统的普及。然而,逐帧重建存在时间不一致问题,并会丢失用户外观特征。自重演方法虽能生成连贯的3D人脸,却难以忠实还原每帧的瞬时表情与光照变化。为此,本文提出一种新方法,通过融合参考视图的3D先验与每帧输入的动态外观,同时保持身份连贯性与帧级外观真实性。该方法基于编码器架构,仅使用表达条件控制的3D GAN生成的合成数据训练,在室内和野外数据集上均实现当前最优的3D重建精度与时间一致性表现。

原文摘要 · Abstract (English)

Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user's appearance. On the other hand, self-reenactment methods can render coherent 3D portraits by driving a 3D avatar built from a single reference image, but fail to faithfully preserve the user's per-frame appearance (e.g., instantaneous facial expression and lighting). As a result, none of these two frameworks is an ideal solution for democratized 3D telepresence. In this work, we address this dilemma and propose a novel solution that maintains both coherent identity and dynamic per-frame appearance to enable the best possible realism. To this end, we propose a new fusion-based method that takes the best of both worlds by fusing a canonical 3D prior from a reference view with dynamic appearance from per-frame input views, producing temporally stable 3D videos with faithful reconstruction of the user's per-frame appearance. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction and temporal consistency on in-studio and in-the-wild datasets. https://research.nvidia.com/labs/amri/projects/coherent3d

3D重建人脸视频动态融合虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。