用扩散模型生成超大规模逼真人体数据,解决3D标注难问题。
PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
- 结合扩散模型与偏好优化,可控生成带3D网格标注的图像。
- 生成50万+样本,图像质量比渲染数据提升76%。
- 适合训练3D人体估计模型的研究者,尤其关注数据质量与多样性。
由于深度歧义和单目图像中3D几何标注困难,获取用于3D人体网格估计的标注数据集极具挑战。现有数据集要么是真实数据(人工标注3D结构但规模有限),要么是合成数据(由3D引擎渲染,标签精确但缺乏真实感、多样性低且成本高)。本文探索第三种路径:生成式数据。我们提出PoseDreamer,一个利用扩散模型生成大规模合成数据的流水线。该方法结合可控图像生成与直接偏好优化实现控制对齐,采用基于课程学习的困难样本挖掘和多阶段质量过滤。这些组件共同保持了3D标签与生成图像之间的自然对应关系,同时优先处理挑战性样本以最大化数据集效用。使用PoseDreamer,我们生成超过50万张高质量合成样本,在图像质量指标上相比渲染基数据集提升76%。在该数据集上训练的模型性能可媲美或优于在真实世界及传统合成数据上训练的模型。此外,将PoseDreamer与合成数据结合,性能优于真实与合成数据的组合,证明了其互补性。我们将公开完整数据集与生成代码。
原文摘要 · Abstract (English)
Acquiring labeled datasets for 3D human mesh estimation is challenging due to depth ambiguities and the inherent difficulty of annotating 3D geometry from monocular images. Existing datasets are either real, with manually annotated 3D geometry and limited scale, or synthetic, rendered from 3D engines that provide precise labels but suffer from limited photorealism, low diversity, and high production costs. In this work, we explore a third path: generated data. We introduce PoseDreamer, a novel pipeline that leverages diffusion models to generate large-scale synthetic datasets with 3D mesh annotations. Our approach combines controllable image generation with Direct Preference Optimization for control alignment, curriculum-based hard sample mining, and multi-stage quality filtering. Together, these components naturally maintain correspondence between 3D labels and generated images, while prioritizing challenging samples to maximize dataset utility. Using PoseDreamer, we generate more than 500,000 high-quality synthetic samples, achieving a 76% improvement in image-quality metrics compared to rendering-based datasets. Models trained on PoseDreamer achieve performance comparable to or superior to those trained on real-world and traditional synthetic datasets. In addition, combining PoseDreamer with synthetic datasets results in better performance than combining real-world and synthetic datasets, demonstrating the complementary nature of our dataset. We will release the full dataset and generation code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。