arXiv:2509.00403cs.CV2025-09

用多视角生成增强单目人体重建细节,提升新视角表现

DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective

  • 用生成模型从不同视角合成人体动作作为监督信号
  • 在多个数据集上优于现有方法,新视角重建更真实
  • 适合做虚拟人像、数字孪生的开发者参考

我们提出一种新框架,从单目视频重建人体虚拟形象。现有方法难以同时捕捉输入中的精细动态细节并生成合理的新视角内容,主要受限于角色模型表达能力不足和观测数据有限。为此,我们利用先进视频生成模型Human4DiT,从替代视角生成人体运动作为额外监督信号。该方法不仅丰富了未见区域的细节,还有效正则化角色表示,减少伪影。此外,我们引入两种互补策略:通过视频微调注入物理身份以保证动作一致性;采用基于分块的去噪算法实现更高分辨率、更细粒度输出。实验表明,本方法超越近期最先进方法,验证了所提策略的有效性。

原文摘要 · Abstract (English)

We present a novel framework to reconstruct human avatars from monocular videos. Recent approaches have struggled either to capture the fine-grained dynamic details from the input or to generate plausible details at novel viewpoints, which mainly stem from the limited representational capacity of the avatar model and insufficient observational data. To overcome these challenges, we propose to leverage the advanced video generative model, Human4DiT, to generate the human motions from alternative perspective as an additional supervision signal. This approach not only enriches the details in previously unseen regions but also effectively regularizes the avatar representation to mitigate artifacts. Furthermore, we introduce two complementary strategies to enhance video generation: To ensure consistent reproduction of human motion, we inject the physical identity into the model through video fine-tuning. For higher-resolution outputs with finer details, a patch-based denoising algorithm is employed. Experimental results demonstrate that our method outperforms recent state-of-the-art approaches and validate the effectiveness of our proposed strategies.

人体重建视频生成单目视觉虚拟形象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。