arXiv:2411.16668cs.CV2024-11被引 5

用扩散模型提升零样本6自由度物体姿态估计性能

Diffusion Features for Zero-Shot 6DoF Object Pose Estimation

  • 采用扩散模型骨干网络替代传统ViT进行零样本姿态估计
  • 在三个标准数据集上平均召回率最高提升27%
  • 适合对零样本3D姿态估计感兴趣的开发者和研究者

零样本物体姿态估计可在无需特定物体训练的情况下从图像中检索物体姿态。近期方法依赖于视觉基础模型(VFM),这些模型是预训练的通用特征提取器。不同训练数据、网络结构和训练范式导致VFM特性各异,当前主流为自监督视觉变压器(ViT)。本研究评估了潜在扩散模型(LDM)骨干网络在零样本姿态估计中的影响。为在统一基础上比较两类模型,我们采用并改进了一种近期方法,提出一种基于模板的多阶段零样本姿态估计方法,使用LDM。在三个标准物体6自由度姿态估计数据集上进行了实证评估。实验表明,该方法相比ViT基线平均召回率提升高达27%。源代码已公开:https://github.com/BvG1993/DZOP。

原文摘要 · Abstract (English)

Zero-shot object pose estimation enables the retrieval of object poses from images without necessitating object-specific training. In recent approaches this is facilitated by vision foundation models (VFM), which are pre-trained models that are effectively general-purpose feature extractors. The characteristics exhibited by these VFMs vary depending on the training data, network architecture, and training paradigm. The prevailing choice in this field are self-supervised Vision Transformers (ViT). This study assesses the influence of Latent Diffusion Model (LDM) backbones on zero-shot pose estimation. In order to facilitate a comparison between the two families of models on a common ground we adopt and modify a recent approach. Therefore, a template-based multi-staged method for estimating poses in a zero-shot fashion using LDMs is presented. The efficacy of the proposed approach is empirically evaluated on three standard datasets for object-specific 6DoF pose estimation. The experiments demonstrate an Average Recall improvement of up to 27% over the ViT baseline. The source code is available at: https://github.com/BvG1993/DZOP.

姿态估计扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。