仅用一张图生成可动3D人像,还能真实模拟衣物变形。
PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image
- 用扩散模型从单图生成多姿态视频,再优化3D avatar
- 在多种姿势下保持形象真实、细节清晰
- 适合想快速制作个性化3D角色的创作者
现有两种构建可动画人体形象的方法:基于3D的方法需大量姿态丰富的视频来建模非刚性形变(如衣物褶皱),但日常难以获取;基于扩散模型的方法虽能从海量野外视频中学习姿态驱动变形,却常导致身份失真或姿态依赖的身份混淆。本文提出PERSONA框架,仅凭一张图像即可生成带有姿态驱动变形的个性化3D人体形象。该方法利用扩散模型从输入图像生成多姿态视频,并以此优化3D avatar。为确保跨姿态下的高保真度与清晰渲染,引入平衡采样策略,通过过采样输入图像缓解扩散生成视频中的身份偏移;采用几何加权优化,优先保证几何约束而非图像损失,有效维持复杂姿态下的渲染质量。
原文摘要 · Abstract (English)
Two major approaches exist for creating animatable human avatars. The first, a 3D-based approach, optimizes a NeRF- or 3DGS-based avatar from videos of a single person, achieving personalization through a disentangled identity representation. However, modeling pose-driven deformations, such as non-rigid cloth deformations, requires numerous pose-rich videos, which are costly and impractical to capture in daily life. The second, a diffusion-based approach, learns pose-driven deformations from large-scale in-the-wild videos but struggles with identity preservation and pose-dependent identity entanglement. We present PERSONA, a framework that combines the strengths of both approaches to obtain a personalized 3D human avatar with pose-driven deformations from a single image. PERSONA leverages a diffusion-based approach to generate pose-rich videos from the input image and optimizes a 3D avatar based on them. To ensure high authenticity and sharp renderings across diverse poses, we introduce balanced sampling and geometry-weighted optimization. Balanced sampling oversamples the input image to mitigate identity shifts in diffusion-generated training videos. Geometry-weighted optimization prioritizes geometry constraints over image loss, preserving rendering quality in diverse poses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。