仅用普通摄像头实现实时3D视频通话,还原真实人物动态视角。
VoluMe -- Authentic 3D Video Calls from Live Gaussian Splat Prediction
- 基于单个2D视频帧独立预测3D高斯点云,保证每帧视觉真实性。
- 在真实视频序列上实现高质量且时间稳定的3D重建,优于现有方法。
- 无需特殊设备或预训练模型,适合大众化3D视频会议应用。
虚拟3D会议有望提升远程协作的共在感与参与度,但实时构建人物3D表示仍具挑战。现有方法依赖复杂硬件、固定外观录入或反演预训练生成模型,难以适配日常视频会议。本文提出首个从单个2D网络摄像头流实时预测3D高斯点云的方法,使3D表示既实时、逼真,又忠实于输入视频(称为“真实性”)。通过独立条件化每一视频帧,重建结果能从捕获视角准确复现原视频,并泛化至新视角。此外,引入稳定性损失以确保视频序列中重建的时序一致性。实验表明,该方法在视觉质量与稳定性指标上达到当前最优水平。我们在仅使用标准2D相机和显示器的情况下,成功演示了实时一对一3D会议。这证明本方法可让任何人以低成本、高真实感的方式进行体积化视频通信。
原文摘要 · Abstract (English)
Virtual 3D meetings offer the potential to enhance copresence, increase engagement and thus improve effectiveness of remote meetings compared to standard 2D video calls. However, representing people in 3D meetings remains a challenge; existing solutions achieve high quality by using complex hardware, making use of fixed appearance via enrolment, or by inverting a pre-trained generative model. These approaches lead to constraints that are unwelcome and ill-fitting for videoconferencing applications. We present the first method to predict 3D Gaussian reconstructions in real time from a single 2D webcam feed, where the 3D representation is not only live and realistic, but also authentic to the input video. By conditioning the 3D representation on each video frame independently, our reconstruction faithfully recreates the input video from the captured viewpoint (a property we call authenticity), while generalizing realistically to novel viewpoints. Additionally, we introduce a stability loss to obtain reconstructions that are temporally stable on video sequences. We show that our method delivers state-of-the-art accuracy in visual quality and stability metrics compared to existing methods, and demonstrate our approach in live one-to-one 3D meetings using only a standard 2D camera and display. This demonstrates that our approach can allow anyone to communicate volumetrically, via a method for 3D videoconferencing that is not only highly accessible, but also realistic and authentic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。