提升无类别姿态估计的匹配精度,解决语义模糊与细节差异问题
CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation
- 引入分层跨模态交互与双流特征优化,融合类别与实例级信息
- 在MP-100数据集上显著超越现有方法,各主干网络均表现更优
- 适合关注细粒度姿态匹配与跨类别泛化能力的研究者
当前无类别姿态估计(CAPE)多采用固定文本关键点描述作为语义先验,虽提升了鲁棒性与灵活性,但静态关节嵌入存在两大固有缺陷:(1)多义性导致跨类别匹配歧义(如“腿”在人与家具中视觉差异大);(2)对类内细粒度差异(如白猫卧姿与黑猫站姿的体态、毛发差异)判别力不足。为此,本文提出新框架,通过层次化跨模态交互与双流特征精炼,将文本描述与具体图像中的类别级和实例级线索融入关节嵌入。在MP-100数据集上的实验表明,无论使用何种网络主干,CapeNext均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。