解决单目视频中人体遮挡导致的渲染失真问题。
Uncertainty-Aware 4D Gaussian Splatting for Monocular Occluded Human Rendering
- 引入概率形变网络与联合光栅化,生成像素级不确定性图
- 在ZJU-MoCap和OcMotion上实现最佳渲染精度与鲁棒性
- 适合关注动态人体重建与不确定性建模的研究者
从单目视频高保真渲染动态人体通常在遮挡情况下严重退化。现有方法要么通过生成模型幻化缺失内容,导致严重时间闪烁;要么采用刚性几何启发式,无法捕捉多样外观。本文将任务重述为异方差观测噪声下的最大后验估计问题,提出U-4DGS框架,融合概率形变网络与联合光栅化管线,生成像素对齐的不确定性图,作为自适应梯度调节器,自动抑制不可靠观测引起的伪影。此外,为防止缺乏可靠视觉线索区域的几何漂移,引入置信度感知正则化,利用学习到的不确定性选择性传播时空有效性。在ZJU-MoCap和OcMotion数据集上的大量实验表明,U-4DGS实现了最先进的渲染保真度与鲁棒性。
原文摘要 · Abstract (English)
High-fidelity rendering of dynamic humans from monocular videos typically degrades catastrophically under occlusions. Existing solutions incorporate external priors-either hallucinating missing content via generative models, which induces severe temporal flickering, or imposing rigid geometric heuristics that fail to capture diverse appearances. To this end, we reformulate the task as a Maximum A Posteriori estimation problem under heteroscedastic observation noise. In this paper, we propose U-4DGS, a framework integrating a Probabilistic Deformation Network and a Joint Rasterization pipeline. This architecture renders pixel-aligned uncertainty maps that act as an adaptive gradient modulator, automatically attenuating artifacts from unreliable observations. Furthermore, to prevent geometric drift in regions lacking reliable visual cues, we enforce Confidence-Aware Regularizations, which leverage the learned uncertainty to selectively propagate spatial-temporal validity. Extensive experiments on the ZJU-MoCap and OcMotion datasets demonstrate that U-4DGS achieves state-of-the-art rendering fidelity and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。