arXiv:2412.03011cs.CVcs.AI2024-12

用单视图模型迁移身体与面部特征,实现高质量多视角人体生成

Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations

  • 基于单视图预训练模型迁移多视角身体表征
  • 融合多模态面部特征提升细节还原能力
  • 在基准数据集上超越现有方法,生成更真实的人体图像

从单视角生成多视角人体图像是一项复杂而重要的挑战。尽管多视角物体生成在扩散模型下取得显著进展,但人体的新视角合成仍受限于3D人体数据集的稀缺。因此,许多现有模型难以生成逼真的人体形态或准确捕捉精细面部细节。为此,我们提出一种创新框架,利用迁移的身体与面部表征进行多视角人体合成。具体而言,我们使用在大规模人体数据集上预训练的单视图模型,构建多视角身体表征,旨在将单视图模型的2D知识扩展至多视角扩散模型。此外,为增强模型的细节恢复能力,我们将迁移的多模态面部特征集成到训练好的人体扩散模型中。在基准数据集上的实验评估表明,该方法优于当前最先进方法,在多视角人体合成任务中表现卓越。

原文摘要 · Abstract (English)

Generating multi-view human images from a single view is a complex and significant challenge. Although recent advancements in multi-view object generation have shown impressive results with diffusion models, novel view synthesis for humans remains constrained by the limited availability of 3D human datasets. Consequently, many existing models struggle to produce realistic human body shapes or capture fine-grained facial details accurately. To address these issues, we propose an innovative framework that leverages transferred body and facial representations for multi-view human synthesis. Specifically, we use a single-view model pretrained on a large-scale human dataset to develop a multi-view body representation, aiming to extend the 2D knowledge of the single-view model to a multi-view diffusion model. Additionally, to enhance the model's detail restoration capability, we integrate transferred multimodal facial features into our trained human diffusion model. Experimental evaluations on benchmark datasets demonstrate that our approach outperforms the current state-of-the-art methods, achieving superior performance in multi-view human synthesis.

人体生成扩散模型多视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。