arXiv:2605.23555cs.CV2026-05

用三模块框架增强单目视频生成3D人像的细节与质量

Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos

论文配图:Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos
图 1 · 摘自论文原文
  • 生成模块通过姿态和相机扰动创造新样本
  • 精修模块用扩散模型提升生成数据质量
  • 评估模块筛选一致性样本,适合细节重建场景

本文针对从单目视频重建逼真且可动画化的3D人体虚拟形象这一挑战。现有方法依赖个体优化结合通用人体先验,但在训练帧数有限时难以捕捉细微特征。为缓解数据稀缺问题,提出TrioMan——一个系统性的三模块数据增强框架。该框架包含三个协同组件:生成器通过在姿态和相机参数上施加高斯扰动,生成多样化未见样本;精修器利用纹理与几何线索引导的一步扩散模型,提升生成数据质量;评估器采用双分支注意力相似性评估,筛选出主体一致的样本。在X-Humans和NeuMan基准上的实验表明,TrioMan优于当前最优方法。

原文摘要 · Abstract (English)

This paper addresses the challenge of reconstructing photorealistic and animatable 3D human avatars from monocular videos. While existing methods rely on combining per-subject optimization with generic human priors, they often fail to capture fine-grained details when training frames are limited. To mitigate this data scarcity, we propose TrioMan, a systematic tri-module framework for augmented 3D avatar learning. Our approach comprises three synergistic components. The Generator creates diverse unseen samples by imposing Gaussian perturbations on pose and camera. The Refiner improves the quality of generated data through one-step diffusion guided by texture and geometry cues. The Examiner selects subject-consistent samples using a dual-branch attention-based similarity evaluation. Experiments on the X-Humans and NeuMan benchmarks show that TrioMan outperforms state-of-the-art methods.

3D人体重建数据增强扩散模型单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。