用可控生成技术融合真实图像与精确人体参数,提升3D人体姿态估计精度。
Toward Human Understanding with Controllable Synthesis
- 结合生成模型与传统图形学,通过身体参数控制生成图像。
- 生成图像越真实,越易偏离真实人体参数,导致模型性能下降。
- 提出新数据集Gen-B,实现高逼真度与精准标注的平衡,适合训练人体姿态网络。
训练鲁棒的3D人体姿态与形状(HPS)估计模型需要多样且带精确标注的训练图像。虽然BEDLAM展示了传统程序化图形生成此类数据的潜力,但其图像明显为合成图像。而生成式图像模型虽能生成高度真实的图像,却缺乏真实标注。将两者结合看似直接:用生成模型并以人体标注作为控制信号。然而我们发现,生成图像越真实,越偏离真实标注,使其不适合作为训练和评估数据。增强如服装、面部表情等细节会引入细微但显著的偏差,可能误导训练。我们通过实验证明,这种偏差会导致使用生成图像训练的HPS网络性能下降。为此,我们设计了一种可控合成方法,在图像真实感与标注精确性之间取得平衡。基于此构建了生成式BEDLAM(Gen-B)数据集,在保持真实标注准确性的前提下显著提升图像真实感。我们通过多种噪声调控策略进行广泛实验,首次证明生成模型可由传统图形方法控制,生成有助于提升HPS方法精度的训练数据。
原文摘要 · Abstract (English)
Training methods to perform robust 3D human pose and shape (HPS) estimation requires diverse training images with accurate ground truth. While BEDLAM demonstrates the potential of traditional procedural graphics to generate such data, the training images are clearly synthetic. In contrast, generative image models produce highly realistic images but without ground truth. Putting these methods together seems straightforward: use a generative model with the body ground truth as controlling signal. However, we find that, the more realistic the generated images, the more they deviate from the ground truth, making them inappropriate for training and evaluation. Enhancements of realistic details, such as clothing and facial expressions, can lead to subtle yet significant deviations from the ground truth, potentially misleading training models. We empirically verify that this misalignment causes the accuracy of HPS networks to decline when trained with generated images. To address this, we design a controllable synthesis method that effectively balances image realism with precise ground truth. We use this to create the Generative BEDLAM (Gen-B) dataset, which improves the realism of the existing synthetic BEDLAM dataset while preserving ground truth accuracy. We perform extensive experiments, with various noise-conditioning strategies, to evaluate the tradeoff between visual realism and HPS accuracy. We show, for the first time, that generative image models can be controlled by traditional graphics methods to produce training data that increases the accuracy of HPS methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。