用概率提示学习提升跨物种动物姿态估计的泛化能力
Probabilistic Prompt Distribution Learning for Animal Pose Estimation
- 设计可学习的概率提示,增强文本描述多样性
- 在多物种基准上实现监督与零样本下的领先性能
- 适合处理长尾分布和未见物种的姿态估计任务
多物种动物姿态估计因视觉差异大、不确定性高而极具挑战。本文针对视觉-语言预训练模型(如CLIP)提出高效提示学习方法,解决跨物种泛化难题。核心在于提示设计、概率提示建模与跨模态适配,使提示能弥补跨模态信息缺失,并有效应对不平衡数据分布下的大尺度数据变异。我们提出一种新颖的概率提示方法,充分挖掘文本描述,缓解长尾分布带来的多样性问题,提升提示对未见类别实例的适应性。具体地,引入一组可学习提示,并通过多样性损失保持提示间差异性,以表征多样图像属性;采样多样化的文本概率表示作为姿态估计引导。随后,在空间层面探索三种跨模态融合策略,减轻视觉不确定性影响。在多个多物种动物姿态基准上的大量实验表明,该方法在监督与零样本设置下均达到当前最优性能。代码已开源:https://github.com/Raojiyong/PPAP。
原文摘要 · Abstract (English)
Multi-species animal pose estimation has emerged as a challenging yet critical task, hindered by substantial visual diversity and uncertainty. This paper challenges the problem by efficient prompt learning for Vision-Language Pretrained (VLP) models, \textit{e.g.} CLIP, aiming to resolve the cross-species generalization problem. At the core of the solution lies in the prompt designing, probabilistic prompt modeling and cross-modal adaptation, thereby enabling prompts to compensate for cross-modal information and effectively overcome large data variances under unbalanced data distribution. To this end, we propose a novel probabilistic prompting approach to fully explore textual descriptions, which could alleviate the diversity issues caused by long-tail property and increase the adaptability of prompts on unseen category instance. Specifically, we first introduce a set of learnable prompts and propose a diversity loss to maintain distinctiveness among prompts, thus representing diverse image attributes. Diverse textual probabilistic representations are sampled and used as the guidance for the pose estimation. Subsequently, we explore three different cross-modal fusion strategies at spatial level to alleviate the adverse impacts of visual uncertainty. Extensive experiments on multi-species animal pose benchmarks show that our method achieves the state-of-the-art performance under both supervised and zero-shot settings. The code is available at https://github.com/Raojiyong/PPAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。