arXiv:2411.11363cs.CV2024-11TPAMI被引 25

无需微调,实时生成高分辨率人体场景图像

GPS-Gaussian+: Generalizable Pixel-wise 3D Gaussian Splatting for Real-Time Human-Scene Rendering from Sparse Views

  • 用参数图直接回归3D高斯属性,实现零优化新视角合成
  • 在稀疏视角下仍保持高帧率,渲染速度超越现有方法
  • 适合交互式应用,尤其适合需快速生成的虚拟角色场景

可微分渲染技术在人物自由视角视频合成方面展现出良好效果。然而,传统方法如高斯点阵或神经隐式渲染通常需要针对每个主体进行优化,难以满足交互应用中的实时渲染需求。本文提出一种通用的高斯点阵方法,在稀疏视角条件下实现高分辨率图像的实时渲染。通过在源视图上定义高斯参数图,并直接回归高斯属性以实现即时新视角合成,无需任何微调或优化。模型在仅含人体或人体-场景数据上训练,联合深度估计模块将2D参数图提升至3D空间。整个框架完全可微,支持深度与渲染双重监督,或仅渲染监督。进一步引入正则化项和对极注意力机制,增强两源视图间的几何一致性,尤其在缺乏深度监督时表现更优。多个数据集上的实验表明,该方法在性能上优于当前最优方法,同时具备卓越的渲染速度。

原文摘要 · Abstract (English)

Differentiable rendering techniques have recently shown promising results for free-viewpoint video synthesis of characters. However, such methods, either Gaussian Splatting or neural implicit rendering, typically necessitate per-subject optimization which does not meet the requirement of real-time rendering in an interactive application. We propose a generalizable Gaussian Splatting approach for high-resolution image rendering under a sparse-view camera setting. To this end, we introduce Gaussian parameter maps defined on the source views and directly regress Gaussian properties for instant novel view synthesis without any fine-tuning or optimization. We train our Gaussian parameter regression module on human-only data or human-scene data, jointly with a depth estimation module to lift 2D parameter maps to 3D space. The proposed framework is fully differentiable with both depth and rendering supervision or with only rendering supervision. We further introduce a regularization term and an epipolar attention mechanism to preserve geometry consistency between two source views, especially when neglecting depth supervision. Experiments on several datasets demonstrate that our method outperforms state-of-the-art methods while achieving an exceeding rendering speed.

3D重建高斯点阵实时渲染稀疏视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。