无需预设骨架,用图像支持样本自动生成关键点结构,提升任意物体姿态估计精度
GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose Estimation

- 从支持图像中自适应推断关键点关系,构建可迭代优化的软连接结构
- 在MP-100数据集上1-shot和5-shot设置下均显著超越图结构基线
- 适合需要快速适配新类别的少样本姿态估计场景
类别无关姿态估计(CAPE)旨在仅通过少量标注的支持样例,对任意类别查询图像中的关键点进行定位。现有方法要么将关键点视为孤立实体,要么依赖人工定义的骨骼先验,后者标注成本高且在不同类别间缺乏灵活性,难以捕捉实例级结构线索,限制了像素级定位精度。为此,我们提出GenCape,一种基于生成的CAPE框架,仅通过图像支持输入即可推断关键点关系,无需额外文本描述或预定义骨架。框架包含两个核心组件:迭代式结构感知变分自编码器(i-SVAE)和组合图迁移模块(CGT)。前者通过变分推理从支持特征中推断出软的、实例特定的邻接矩阵,并分层嵌入图变换解码器以逐步优化结构先验;后者通过贝叶斯融合与注意力重加权,自适应聚合多个潜在图结构,形成查询感知结构,增强对视觉不确定性与支持样本偏差的鲁棒性。该结构感知设计促进了关键点间的有效信息传播,并在具有多样拓扑的关键点结构间实现语义对齐。在MP-100数据集上的实验表明,本方法在1-shot和5-shot设置下均显著优于图结构基线,同时保持与文本支持方法相当的性能。
原文摘要 · Abstract (English)
Category-agnostic pose estimation (CAPE) aims to localize keypoints on query images from arbitrary categories, using only a few annotated support examples for guidance. Recent approaches either treat keypoints as isolated entities or rely on manually defined skeleton priors, which are costly to annotate and inherently inflexible across diverse categories. Such oversimplification limits the model's capacity to capture instance-wise structural cues critical for accurate pixel-level localization. To overcome these limitations, we propose GenCape, a Generative-based framework for CAPE that infers keypoint relationships solely from image-based support inputs, without additional textual descriptions or predefined skeletons. Our framework consists of two principal components: an iterative Structure-aware Variational Autoencoder (i-SVAE) and a Compositional Graph Transfer (CGT) module. The former infers soft, instance-specific adjacency matrices from support features through variational inference, embedded layer-wise into the Graph Transformer Decoder for progressive structural priors refinement. The latter adaptively aggregates multiple latent graphs into a query-aware structure via Bayesian fusion and attention-based reweighting, enhancing resilience to visual uncertainty and support-induced bias. This structure-aware design facilitates effective message propagation among keypoints and promotes semantic alignment across object categories with diverse keypoint topologies. Experimental results on the MP-100 dataset show that our method achieves substantial gains over graph-support baselines under both 1- and 5-shot settings, while maintaining competitive performance against text-support counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。