arXiv:2410.18355cs.CVcs.GR2024-10CVPR被引 31

实时重光照真人视频,支持任意视角和灯光变化。

Real-time 3D-aware Portrait Video Relighting

  • 用双编码器快速提取肤色与光影三平面,实现3D解耦表示。
  • 在消费级硬件上达32.98帧/秒,重建质量与时间一致性领先。
  • 适合需要交互式视频重光照的会议、影视制作场景。

在视频会议等应用中,对说话人脸在自定义光照与视角下生成逼真视频具有重要意义。然而,现有方法或计算耗时,或无法调整视角。本文提出首个基于神经辐射场(NeRF)的实时3D感知人脸视频重光照方法。给定输入的人脸视频,该方法可生成在新视角与新光照条件下的逼真说话人脸,具备解耦的3D表示。具体地,通过快速双编码器为每帧推断出肤色三平面与目标光照下的阴影三平面,并利用时序一致性网络保证过渡平滑、减少闪烁。方法在消费级硬件上达到32.98帧/秒,且在重建质量、光照误差、光照不稳定性、时序一致性和推理速度上均达到当前最优。实验展示了其在多种真实场景视频上的有效性与交互性。

原文摘要 · Abstract (English)

Synthesizing realistic videos of talking faces under custom lighting conditions and viewing angles benefits various downstream applications like video conferencing. However, most existing relighting methods are either time-consuming or unable to adjust the viewpoints. In this paper, we present the first real-time 3D-aware method for relighting in-the-wild videos of talking faces based on Neural Radiance Fields (NeRF). Given an input portrait video, our method can synthesize talking faces under both novel views and novel lighting conditions with a photo-realistic and disentangled 3D representation. Specifically, we infer an albedo tri-plane, as well as a shading tri-plane based on a desired lighting condition for each video frame with fast dual-encoders. We also leverage a temporal consistency network to ensure smooth transitions and reduce flickering artifacts. Our method runs at 32.98 fps on consumer-level hardware and achieves state-of-the-art results in terms of reconstruction quality, lighting error, lighting instability, temporal consistency and inference speed. We demonstrate the effectiveness and interactivity of our method on various portrait videos with diverse lighting and viewing conditions.

视频重光照实时生成3D人脸NeRF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。