用轻量模型实现低码率下高保真3D人脸通话视频压缩。
Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference
- 结合参数化人脸模型与高斯渲染,只传关键面部数据。
- 在极低码率下仍保持高质量人脸还原,性能优于传统方法。
- 适合实时3D视频会议,尤其对带宽敏感的场景
沉浸式互动通信需求推动了3D视频会议的发展,但如何在低码率下实现高保真3D说话人脸表示仍是难题。传统2D视频压缩无法保留精细几何与外观细节,而基于NeRF的隐式神经渲染方法计算开销过大。为此,我们提出一种轻量、高保真、低码率的3D说话人脸压缩框架,融合基于FLAME的参数化建模与3DGS神经渲染。该方法仅实时传输关键面部元数据,借助高斯头模型实现高效重建。同时引入紧凑表示与压缩方案,包括高斯属性压缩和MLP优化,提升传输效率。实验表明,本方法在率失真性能上表现优异,在极低码率下仍能实现高质量人脸渲染,适用于实时3D视频会议应用。
原文摘要 · Abstract (English)
The demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D video compression techniques fail to preserve fine-grained geometric and appearance details, while implicit neural rendering methods like NeRF suffer from prohibitive computational costs. To address these challenges, we propose a lightweight, high-fidelity, low-bitrate 3D talking face compression framework that integrates FLAME-based parametric modeling with 3DGS neural rendering. Our approach transmits only essential facial metadata in real time, enabling efficient reconstruction with a Gaussian-based head model. Additionally, we introduce a compact representation and compression scheme, including Gaussian attribute compression and MLP optimization, to enhance transmission efficiency. Experimental results demonstrate that our method achieves superior rate-distortion performance, delivering high-quality facial rendering at extremely low bitrates, making it well-suited for real-time 3D video conferencing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。