用单图+姿态对模型微调,让3D人体重建更自然。
Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
- 基于姿态与单图配对数据,直接优化生成一致性。
- 在15000个复杂姿态样本上训练,显著提升难姿态还原效果。
- 无需昂贵3D数据,适合研究真实场景下的人体重建。
单视角3D人体重建虽借助多视角扩散模型取得显著进展,但复现的人体常出现不自然姿态,尤其在动态或复杂姿态下更为明显,这归因于现有3D人体数据集姿态多样性不足。为此,本文提出DrPose:一种直接基于姿态的奖励微调算法,可在不依赖昂贵3D人体资产的前提下,对多视角扩散模型进行后训练。DrPose利用仅包含单图与对应姿态的配对数据,通过最大化我们提出的可微分奖励函数PoseScore,优化生成结果与真实姿态的一致性。该方法基于新构建的DrPose15K数据集,其源自已有动作数据集与姿态条件视频生成模型,具有比现有3D数据集更广的姿态分布。我们在标准基准、野外图像及新构建的基准上验证了该方法,重点评估复杂姿态下的表现。结果表明,所有测试中均实现一致的定性和定量提升。
原文摘要 · Abstract (English)
Single-view 3D human reconstruction has achieved remarkable progress through the adoption of multi-view diffusion models, yet the recovered 3D humans often exhibit unnatural poses. This phenomenon becomes pronounced when reconstructing 3D humans with dynamic or challenging poses, which we attribute to the limited scale of available 3D human datasets with diverse poses. To address this limitation, we introduce DrPose, Direct Reward fine-tuning algorithm on Poses, which enables post-training of a multi-view diffusion model on diverse poses without requiring expensive 3D human assets. DrPose trains a model using only human poses paired with single-view images, employing a direct reward fine-tuning to maximize PoseScore, which is our proposed differentiable reward that quantifies consistency between a generated multi-view latent image and a ground-truth human pose. This optimization is conducted on DrPose15K, a novel dataset that was constructed from an existing human motion dataset and a pose-conditioned video generative model. Constructed from abundant human pose sequence data, DrPose15K exhibits a broader pose distribution compared to existing 3D human datasets. We validate our approach through evaluation on conventional benchmark datasets, in-the-wild images, and a newly constructed benchmark, with a particular focus on assessing performance on challenging human poses. Our results demonstrate consistent qualitative and quantitative improvements across all benchmarks. Project page: https://seunguk-do.github.io/drpose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。