arXiv:2503.13025cs.CVcs.AI2025-03ICCV被引 2

从野外2D姿态数据生成多样3D姿态,提升模型泛化能力

PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D Data

  • 通过提取困难姿态并合成运动序列,生成逼真3D训练数据
  • 在真实场景下将多种3D姿态估计器准确率提升最高14%
  • 无需昂贵3D标注,适用于不同模型结构与规模

尽管已有大量研究致力于在不依赖昂贵3D标注的情况下提升3D姿态估计器的泛化能力,现有数据增强方法在面对多样化人体外观和复杂姿态的真实场景时仍表现不佳。本文提出PoseSyn,一种新颖的数据合成框架,可将丰富的野外2D姿态数据转换为多样化的3D姿态图像对。该框架包含两个关键组件:误差提取模块(EEM),用于从2D姿态数据集中识别具有挑战性的姿态;运动合成模块(MSM),用于围绕这些挑战性姿态生成运动序列。随后,通过结合挑战性姿态与外观的人体动画模型生成逼真的3D训练数据,PoseSyn在包含不同背景、遮挡、复杂姿态及多视角场景的真实世界基准上,将多种3D姿态估计器的性能提升最高达14%。大量实验进一步验证,PoseSyn是一种可扩展且高效的方法,无需依赖昂贵3D标注,无论姿态估计器的模型大小或结构如何,均能有效提升其泛化能力。

原文摘要 · Abstract (English)

Despite considerable efforts to enhance the generalization of 3D pose estimators without costly 3D annotations, existing data augmentation methods struggle in real world scenarios with diverse human appearances and complex poses. We propose PoseSyn, a novel data synthesis framework that transforms abundant in the wild 2D pose dataset into diverse 3D pose image pairs. PoseSyn comprises two key components: Error Extraction Module (EEM), which identifies challenging poses from the 2D pose datasets, and Motion Synthesis Module (MSM), which synthesizes motion sequences around the challenging poses. Then, by generating realistic 3D training data via a human animation model aligned with challenging poses and appearances PoseSyn boosts the accuracy of various 3D pose estimators by up to 14% across real world benchmarks including various backgrounds and occlusions, challenging poses, and multi view scenarios. Extensive experiments further confirm that PoseSyn is a scalable and effective approach for improving generalization without relying on expensive 3D annotations, regardless of the pose estimator's model size or design.

3D姿态数据增强姿态合成无标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。