arXiv:2509.04145cs.GRcs.CV2025-09被引 4

用扩散模型生成高保真动态人像,支持跨身份实时渲染。

Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion

  • 通过优化特定人物的UNet网络,捕捉姿态依赖形变
  • 在权重空间训练超扩散模型,实现跨身份生成
  • 支持实时可控渲染,效果优于现有方法

生成逼真人像是一项重要但具挑战的任务。近期基于辐射场渲染的方法在个性化动态人像上实现了前所未有的逼真度和实时性能,但通常仅针对单个人物的多视角视频数据训练,难以跨身份泛化。而基于预训练2D扩散模型的生成方法虽可生成卡通化静态人像并简单通过骨骼动画,但渲染质量较低,无法捕捉衣物褶皱等姿态相关形变。本文提出一种新方法,融合特定人物渲染与扩散生成的优势,实现兼具高保真与真实姿态形变的动态人像生成。采用两阶段流程:首先为每位人物优化一组特定的UNet网络,每个网络代表一个能捕捉复杂姿态形变的动态人像;第二阶段,在已优化的网络权重空间上训练超扩散模型。推理时,该方法生成网络权重,实现对动态人像的实时可控渲染。使用大规模跨身份多视角视频数据集,证明本方法优于当前最先进的人像生成技术。

原文摘要 · Abstract (English)

Creating human avatars is a highly desirable yet challenging task. Recent advancements in radiance field rendering have achieved unprecedented photorealism and real-time performance for personalized dynamic human avatars. However, these approaches are typically limited to person-specific rendering models trained on multi-view video data for a single individual, limiting their ability to generalize across different identities. On the other hand, generative approaches leveraging prior knowledge from pre-trained 2D diffusion models can produce cartoonish, static human avatars, which are animated through simple skeleton-based articulation. Therefore, the avatars generated by these methods suffer from lower rendering quality compared to person-specific rendering methods and fail to capture pose-dependent deformations such as cloth wrinkles. In this paper, we propose a novel approach that unites the strengths of person-specific rendering and diffusion-based generative modeling to enable dynamic human avatar generation with both high photorealism and realistic pose-dependent deformations. Our method follows a two-stage pipeline: first, we optimize a set of person-specific UNets, with each network representing a dynamic human avatar that captures intricate pose-dependent deformations. In the second stage, we train a hyper diffusion model over the optimized network weights. During inference, our method generates network weights for real-time, controllable rendering of dynamic human avatars. Using a large-scale, cross-identity, multi-view video dataset, we demonstrate that our approach outperforms state-of-the-art human avatar generation methods.

人像生成扩散模型动态渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。