用神经雅可比场实现单目视频中人体动画的时序一致重建
JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video

- 通过求解泊松方程预测姿态依赖的雅可比矩阵,建模半刚性形变
- 在真实视频上生成几何一致且时序稳定的虚拟人像,优于现有方法
- 适合需要高保真人体动画的影视与游戏制作场景
复杂动作下的人体虚拟形象生成——如衣物动态——需要同时建模全局与局部形变,但在单目视频设置下仍具挑战。本文提出利用神经雅可比场(NJFs)表示半刚性形变,训练自监督神经网络预测雅可比矩阵以生成姿态相关的形变,通过求解泊松方程实现。针对单目输入带来的遮挡区域与不可见表面问题,引入三项关键设计:约束泊松求解器、基于符号距离的雅可比正则化、以及形变引导的残差流损失,共同抑制边界伪影,恢复腋窝、大腿等高频遮挡区域,并保证运动过程中的时序一致性。在基准数据集与真实场景视频上的实验表明,本方法生成的虚拟人像在时间上稳定且几何一致,性能超越当前最优方法。
原文摘要 · Abstract (English)
Generating realistic human avatars in complex motions--such as clothing dynamics--requires modeling of global and local deformations which remains challenging in monocular settings. We address this problem by leveraging neural Jacobian fields (NJFs) for representing semi-rigid deformations. We train self-supervised neural networks for predicting Jacobian matrices that give the pose-dependent deformations, by solving a Poisson equation. However, monocular input presents several difficulties such as self-occluded regions and invisible surfaces. To address these issues, we introduce three key components: a constrained Poisson solver, signed distance-based Jacobian regularization, and a deformation-guided residual flow loss, which together suppress boundary artifacts, recover frequently occluded regions such as armpits and thighs, and enforce temporal consistency during motion. Experiments on benchmark and in-the-wild videos demonstrate that our method generates temporally stable and geometrically coherent avatars, outperforming state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。