arXiv:2603.28997cs.CV2026-03

单目视频实时捕捉人体动作,动态更新参考空间提升重建质量

GenFusion: Feed-forward Human Performance Capture via Progressive Canonical Space Updates

  • 用渐进更新的通用空间存储历史外观信息
  • 在无直接观测区域仍能生成清晰人体图像
  • 适合需要实时高质量人体重建的应用场景

我们提出一种前馈式人体动作捕捉方法,可从单目RGB视频流中渲染表演者的新视角图像。该任务的核心挑战在于缺乏足够观测,尤其是未见过的区域。假设表演者随时间连续运动,我们利用每个新帧不断更新一个通用空间,累积随时间演化的外观信息,作为当前帧缺失观测时的上下文参考。为有效利用此上下文并尊重实时形变,我们将渲染过程建模为概率回归,解决历史与当前观测之间的冲突,相比确定性回归方法产生更锐利的重建结果。即使在无先前观测的区域,也能实现合理合成。在域内数据集(4D-Dress)和域外数据集(MVHumanNet)上的实验验证了方法的有效性。

原文摘要 · Abstract (English)

We present a feed-forward human performance capture method that renders novel views of a performer from a monocular RGB stream. A key challenge in this setting is the lack of sufficient observations, especially for unseen regions. Assuming the subject moves continuously over time, we take advantage of the fact that more body parts become observable by maintaining a canonical space that is progressively updated with each incoming frame. This canonical space accumulates appearance information over time and serves as a context bank when direct observations are missing in the current live frame. To effectively utilize this context while respecting the deformation of the live state, we formulate the rendering process as probabilistic regression. This resolves conflicts between past and current observations, producing sharper reconstructions than deterministic regression approaches. Furthermore, it enables plausible synthesis even in regions with no prior observations. Experiments on in-domain (4D-Dress) and out-of-distribution (MVHumanNet) datasets demonstrate the effectiveness of our approach.

人体捕捉单目视频生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。