arXiv:2504.00394cs.CV2025-04

用可控生成技术合成高质量动物姿态数据,解决标注稀缺难题。

AP-CAP: Advancing High-Quality Data Synthesis for Animal Pose Estimation via a Controllable Image Generation Pipeline

  • 构建可控制的图像生成流水线,按需生成带目标姿态的动物图像。
  • 提出三策略提升生成质量:多模态融合、动态姿态调整、文本增强。
  • 创建首个混合真实与合成数据的大型基准数据集MPCH,适合行为分析研究者使用。

2D动物姿态估计在动物行为分析和生态研究中至关重要。尽管现有方法取得进展,但高质量数据集稀缺仍是主要瓶颈。为此,本文提出一种可控图像生成流水线AP-CAP,包含支持预期姿态生成的多模态动物图像生成模型。为提升生成数据的质量与多样性,提出三项创新策略:(1)基于多源外观表征融合的图像合成;(2)基于姿态调整的多样化姿态捕捉;(3)基于文本增强的视觉语义理解丰富。利用该模型与策略,构建了首个混合真实与合成数据的MPCH数据集(Modality-Pose-Caption Hybrid),成为当前规模最大、多源异构的动物姿态估计基准库。大量实验表明,该方法显著提升姿态估计算法的性能与泛化能力。

原文摘要 · Abstract (English)

The task of 2D animal pose estimation plays a crucial role in advancing deep learning applications in animal behavior analysis and ecological research. Despite notable progress in some existing approaches, our study reveals that the scarcity of high-quality datasets remains a significant bottleneck, limiting the full potential of current methods. To address this challenge, we propose a novel Controllable Image Generation Pipeline for synthesizing animal pose estimation data, termed AP-CAP. Within this pipeline, we introduce a Multi-Modal Animal Image Generation Model capable of producing images with expected poses. To enhance the quality and diversity of the generated data, we further propose three innovative strategies: (1) Modality-Fusion-Based Animal Image Synthesis Strategy to integrate multi-source appearance representations, (2) Pose-Adjustment-Based Animal Image Synthesis Strategy to dynamically capture diverse pose variations, and (3) Caption-Enhancement-Based Animal Image Synthesis Strategy to enrich visual semantic understanding. Leveraging the proposed model and strategies, we create the MPCH Dataset (Modality-Pose-Caption Hybrid), the first hybrid dataset that innovatively combines synthetic and real data, establishing the largest-scale multi-source heterogeneous benchmark repository for animal pose estimation to date. Extensive experiments demonstrate the superiority of our method in improving both the performance and generalization capability of animal pose estimators.

姿态估计数据合成生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。