通过可控生成增强3D人体姿态估计的域泛化能力
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation

- 构建可控人体姿态生成框架,动态调节姿态、背景和视角
- 在未见场景和数据集上显著提升模型性能,跨域泛化效果明显
- 适合关注真实部署环境下模型鲁棒性的研究者
行人运动具有因果性,其表现受训练与测试数据分布差异导致的域差距强烈影响。针对3D人体姿态估计任务,本文提出一种可控人体姿态生成框架,通过系统性地改变姿态、背景和相机视角,合成多样化的视频数据。该生成增强方法丰富了训练数据,提升了模型的域泛化能力,缓解了现有方法在处理域差异时的局限性。结合室内/真实世界与室外/虚拟数据集,实现跨域数据融合与可控视频生成,构建适用于真实部署场景的增强训练数据。大量实验表明,经增强的数据集显著提升模型在未见场景和数据集上的表现,验证了该方法的有效性。
原文摘要 · Abstract (English)
Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human pose generation framework that synthesizes diverse video data by systematically varying poses, backgrounds, and camera viewpoints. This generative augmentation enriches training datasets, enhances model generalization, and alleviates the limitations of existing methods in handling domain discrepancies. By leveraging both indoor/real-world and outdoor/virtual datasets, we perform cross-domain data fusion and controllable video generation to construct enriched training data, tailored to realistic deployment settings. Extensive experiments show that the augmented datasets significantly improve model performance on unseen scenarios and datasets, validating the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。