用生成先验精修单视角人体图像,实现逼真姿态与视角切换。
Blur2Sharp: Human Novel Pose and View Synthesis with Generative Prior Refinement
- 结合3D神经渲染与扩散模型,分步生成几何一致的多视角图像。
- 在松散衣物和遮挡场景下仍保持清晰细节,优于现有方法。
- 适合需要高保真人体图像生成的研究者或虚拟试衣应用。
构建逼真人体虚拟形象以实现自然姿态变化和视角灵活性仍是计算机视觉与图形学中的核心挑战。现有方法通常产生几何不一致的多视角图像,或牺牲照片真实感,导致在不同视角和复杂动作下输出模糊。为此,我们提出Blur2Sharp,一种融合3D感知神经渲染与扩散模型的新框架,仅需单参考视角即可生成清晰、几何一致的新型视角图像。方法采用双条件架构:首先,通过人体NeRF模型为目标姿态生成几何一致的多视角渲染图,显式编码三维结构引导;随后,利用扩散模型对这些渲染图进行条件化精修,保留细粒度细节与结构保真度。我们进一步通过层次化特征融合,引入参数化SMPL模型提取的纹理、法线与语义先验,同时提升全局一致性与局部细节精度。大量实验表明,Blur2Sharp在新型姿态与视角生成任务中持续超越现有最先进方法,尤其在松散衣物与遮挡等挑战场景下表现优异。
原文摘要 · Abstract (English)
The creation of lifelike human avatars capable of realistic pose variation and viewpoint flexibility remains a fundamental challenge in computer vision and graphics. Current approaches typically yield either geometrically inconsistent multi-view images or sacrifice photorealism, resulting in blurry outputs under diverse viewing angles and complex motions. To address these issues, we propose Blur2Sharp, a novel framework integrating 3D-aware neural rendering and diffusion models to generate sharp, geometrically consistent novel-view images from only a single reference view. Our method employs a dual-conditioning architecture: initially, a Human NeRF model generates geometrically coherent multi-view renderings for target poses, explicitly encoding 3D structural guidance. Subsequently, a diffusion model conditioned on these renderings refines the generated images, preserving fine-grained details and structural fidelity. We further enhance visual quality through hierarchical feature fusion, incorporating texture, normal, and semantic priors extracted from parametric SMPL models to simultaneously improve global coherence and local detail accuracy. Extensive experiments demonstrate that Blur2Sharp consistently surpasses state-of-the-art techniques in both novel pose and view generation tasks, particularly excelling under challenging scenarios involving loose clothing and occlusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。