用双未投影纹理实现稀疏视角下实时高保真人体渲染
Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures
- 分离几何形变与外观合成,提升渲染鲁棒性
- 4K分辨率下实现实时渲染,显著优于现有方法
- 适合需要高质量动态人体重建的应用场景
从稀疏视角的RGB视频实现实时自由视角人体渲染极具挑战,因传感器数量有限且时间预算紧张。现有方法多依赖2D CNN在纹理空间学习渲染基元,但或联合学习几何与外观,或完全忽略稀疏图像信息进行几何估计,严重损害视觉质量及对未见姿态的鲁棒性。为此,本文提出双未投影纹理(Double Unprojected Textures),核心在于将粗略几何形变估计与外观合成解耦,实现4K级实时、逼真的渲染。具体地,首先引入一种图像条件化的模板形变网络,从首个未投影纹理估计人体模板的粗略形变;该更新后的几何结构用于第二次更精确的纹理反投影,生成伪影更少、与输入视图对齐更好的纹理图,从而促进以高斯斑点表示的细粒度几何与外观的学习。定量与定性实验验证了该方法的有效性与高效性,显著超越其他前沿方法。
原文摘要 · Abstract (English)
Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and appearance, or completely ignore sparse image information for geometry estimation, significantly harming visual quality and robustness to unseen body poses. To address these issues, we present Double Unprojected Textures, which at the core disentangles coarse geometric deformation estimation from appearance synthesis, enabling robust and photorealistic 4K rendering in real-time. Specifically, we first introduce a novel image-conditioned template deformation network, which estimates the coarse deformation of the human template from a first unprojected texture. This updated geometry is then used to apply a second and more accurate texture unprojection. The resulting texture map has fewer artifacts and better alignment with input views, which benefits our learning of finer-level geometry and appearance represented by Gaussian splats. We validate the effectiveness and efficiency of the proposed method in quantitative and qualitative experiments, which significantly surpasses other state-of-the-art methods. Project page: https://vcai.mpi-inf.mpg.de/projects/DUT/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。