用3D结构引导视频扩散模型,实现实时高保真上半身虚拟人生成。
ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
- 结合3D重建提供结构先验,指导实时自回归视频扩散模型渲染。
- 显著减少纹理模糊与动作僵硬,动态连贯性优于主流方法。
- 适合游戏、VR等需要实时高质虚拟人的场景应用。
从单张输入图像生成高保真上半身3D虚拟人仍面临挑战。现有3D生成方法依赖大型重建模型,虽快速稳定,但常出现纹理模糊和动作僵硬;而生成式视频模型虽能合成逼真动态结果,却易出现结构错误和身份漂移。为此,我们提出一种新方法:利用3D重建模型提供稳健的结构与外观先验,指导实时自回归视频扩散模型进行渲染。该流程可实现实时生成高频细节与流畅动态,有效缓解纹理模糊与动作僵硬问题,同时避免视频生成中常见的结构不一致。通过融合3D重建的几何稳定性与视频模型的生成能力,本方法在视觉质量上显著优于当前领先方法,为游戏、虚拟现实等实时应用提供了高效可靠的解决方案。
原文摘要 · Abstract (English)
Generating high-fidelity upper-body 3D avatars from one-shot input image remains a significant challenge. Current 3D avatar generation methods, which rely on large reconstruction models, are fast and capable of producing stable body structures, but they often suffer from artifacts such as blurry textures and stiff, unnatural motion. In contrast, generative video models show promising performance by synthesizing photorealistic and dynamic results, but they frequently struggle with unstable behavior, including body structural errors and identity drift. To address these limitations, we propose a novel approach that combines the strengths of both paradigms. Our framework employs a 3D reconstruction model to provide robust structural and appearance priors, which in turn guides a real-time autoregressive video diffusion model for rendering. This process enables the model to synthesize high-frequency, photorealistic details and fluid dynamics in real time, effectively reducing texture blur and motion stiffness while preventing the structural inconsistencies common in video generation methods. By uniting the geometric stability of 3D reconstruction with the generative capabilities of video models, our method produces high-fidelity digital avatars with realistic appearance and dynamic, temporally coherent motion. Experiments demonstrate that our approach significantly reduces artifacts and achieves substantial improvements in visual quality over leading methods, providing a robust and efficient solution for real-time applications such as gaming and virtual reality. Project page: https://lhyfst.github.io/visa
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。