通过软硬件协同优化,实现虚拟现实中的高保真表情动画实时渲染。
ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
- 针对编码人脸模型设计低精度量化方法,保持图像质量
- 在4比特下提升视觉评分0.39,延迟降低3.36倍
- 适合资源受限的头戴式设备,支持每秒100帧实时渲染
基于深度学习的高保真编码人脸(PCA)在虚拟现实中用于实现沉浸式交互,但其计算需求高,在头戴式显示设备等资源受限场景中难以实现实时推理。为此,我们提出一种面向编码人脸模型的后训练量化(PTQ)方法,可在不损失输出质量的前提下实现低精度执行。同时,设计了一款可集成于系统芯片的专用硬件加速器,进一步提升处理效率。基于此,构建了端到端的全栈优化框架ESCA,显著提升边缘VR平台上的PCA推理性能。实验表明,ESCA在FovVideoVDP质量评分上比最佳4比特基线提升0.39,延迟降低达3.36倍,并在端到端测试中维持100帧每秒的渲染速率,满足实时虚拟现实需求。该工作证明了在资源受限设备上部署高保真编码人脸的可行性,为更沉浸、便携的虚拟现实体验开辟新路径。
原文摘要 · Abstract (English)
Photorealistic Codec Avatars (PCA), which generate high-fidelity human face renderings, are increasingly being used in Virtual Reality (VR) environments to enable immersive communication and interaction through deep learning-based generative models. However, these models impose significant computational demands, making real-time inference challenging on resource-constrained VR devices such as head-mounted displays, where latency and power efficiency are critical. To address this challenge, we propose an efficient post-training quantization (PTQ) method tailored for Codec Avatar models, enabling low-precision execution without compromising output quality. In addition, we design a custom hardware accelerator that can be integrated into the system-on-chip of VR devices to further enhance processing efficiency. Building on these components, we introduce ESCA, a full-stack optimization framework that accelerates PCA inference on edge VR platforms. Experimental results demonstrate that ESCA boosts FovVideoVDP quality scores by up to $+0.39$ over the best 4-bit baseline, delivers up to $3.36\times$ latency reduction, and sustains a rendering rate of 100 frames per second in end-to-end tests, satisfying real-time VR requirements. These results demonstrate the feasibility of deploying high-fidelity codec avatars on resource-constrained devices, opening the door to more immersive and portable VR experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。