用单视角视频给静态神经点云加骨骼动画,避免关节处的缝隙和突起。
RigPAPR: Rig-Based Animation of Static Neural Point Clouds from a Fixed-Viewpoint Video

- 通过动态重组成像避免点云固定形状导致的关节畸变。
- 在新视角下比网格和高斯点云方法提升3dB以上PSNR。
- 无需网格、模板或额外修正,直接从视频自动绑定骨骼。
静态神经点云重建能以高保真度捕捉人物姿态。给定此类重建,我们旨在将其动画化,以匹配单视角驱动视频(无论是真实拍摄还是图像生成视频),并恢复一个可重新绑定的带骨骼3D资产。现有方法通过直接线性混合皮肤(LBS)变形高斯斑点或网格代理,但在关节活动时仍易产生边界伪影,即使使用逐原语修正也难以避免。我们发现伪影源于表示方式:每个斑点在标准姿态中独立校准形状,与其他斑点拼接。在刚性LBS下,斑点随骨骼移动但无法弯曲,导致关节边界处出现间隙或突起。相比之下,邻近注意力点渲染(PAPR)不携带每个原语的固定形状;每个像素在渲染时由变形原语位置重新合成,使表面能自然随动作重组。我们提出RigPAPR,可自动为静态PAPR点云绑定骨骼,并仅凭单一固定视角视频驱动其进行直接LBS动画,无需网格代理、姿态依赖修正或类别模板。在合成对象上,RigPAPR在监督视图上达到最强基线性能,在新视角下优于基于网格和高斯点云的方法,提升超过3dB PSNR,且合成与真实对象的关节边界渲染更清晰。
原文摘要 · Abstract (English)
Static neural point reconstructions capture a subject at high fidelity from posed images. Given such a reconstruction, we aim to animate it to follow a monocular fixed-viewpoint driving video of the subject, whether captured or produced by image-to-video (I2V) generation, and to recover a rigged, re-posable 3D asset. Existing methods deform Gaussian splats through direct linear blend skinning (LBS) or mesh proxies, both of which are prone to joint-boundary artifacts under articulation, even with per-primitive corrections. We trace the artifact to the representation: each splat carries an individual shape calibrated in the canonical pose to tile with its neighbours. Under rigid LBS, each splat moves with its bone but cannot bend, so the canonical tiling breaks at joint boundaries into gaps and spikes. Proximity attention point rendering (PAPR) instead carries no per-primitive shape; each pixel is recomposed at render time from the deformed primitives' positions, so the surface re-forms naturally with the articulation. We present RigPAPR, which auto-rigs a static PAPR cloud and drives it under direct LBS from a single fixed-viewpoint video, without mesh proxy, pose-dependent correction, or category template. On synthetic subjects, RigPAPR matches the strongest baseline at the supervised view and exceeds mesh-based and Gaussian-splatting baselines at novel views by 3+dB PSNR, with cleaner joint-boundary renderings of both synthetic and real subjects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。