通过多视角语义一致性,提升稀疏视图下人体3D高斯点的定位精度。
Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency

- 利用跨视角注意力重校准同部位特征,统一多视角表示
- 在多个基准数据集上显著提升稀疏视图下人体渲染质量
- 特别适合复杂肢体动作与视角重叠少的场景
近期,基于稀疏视图输入的可泛化人体高斯点云渲染受到广泛关注。现有方法多依赖显式几何约束或预定义结构表示来精确定位3D高斯点,尽管取得显著进展,但仍因人体复杂动作和视图间重叠有限,导致多视角特征表示不一致。为此,本文提出一种新方法:通过预测深度图将各视角编码的潜在嵌入反投影至共享3D空间,并基于跨视角注意力重新校准同一身体部位的特征。该机制有效缓解了高纹理区域及遮挡部位的空间模糊问题,实现3D高斯点的精准定位。在多个基准数据集上的实验表明,该方法能高效提升稀疏视图下可泛化人体高斯点云渲染性能。
原文摘要 · Abstract (English)
Recently, generalizable human Gaussian splatting from sparse-view inputs has been actively studied for the photorealistic human rendering. Most existing methods rely on explicit geometric constraints or predefined structural representations to accurately position 3D Gaussians. Although these approaches have shown the remarkable progress in this field, they still suffer from inconsistent feature representations across multi-view inputs due to complex articulations of the human body and limited overlaps between different views. To address this problem, we propose a novel method to accurately localize 3D Gaussians and ultimately improve the quality of human rendering. The key idea is to unproject latent embeddings encoded from each viewpoint into a shared 3D space through predicted depth maps and recalibrate them belonging to the same body part based on cross-view attention. This helps the model resolve the spatial ambiguity occurring in highly textured regions as well as occluded body parts, thus leading to the accurate localization of 3D Gaussians. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of generalizable human Gaussian splatting from sparse-view inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。