arXiv:2501.01993cs.CVcs.LG2025-01被引 1

用新卷积和注意力机制,提升单目图像6D姿态估计精度

A Novel Convolution and Attention Mechanism-based Model for 6D Object Pose Estimation

  • 构建图像特征图的图结构,显式建模空间关系
  • 在LINEMOD等数据集上优于传统方法,尤其在遮挡场景表现更优
  • 适合做高精度物体姿态估计的研究者和工业应用开发者

本文提出PoseLecTr,一种基于图的编码器-解码器框架,结合新型勒让德卷积与注意力机制,实现从单目RGB图像中进行六自由度(6-DOF)物体姿态估计。传统学习方法主要依赖网格结构卷积,难以有效建模图像特征间的高阶及长程依赖关系,尤其在杂乱或遮挡场景下表现受限。PoseLecTr通过将图像特征构造成图表示,以图连接性显式建模空间关系。提出的框架引入勒让德卷积层提升图卷积的数值稳定性,并结合空间注意力与自注意力蒸馏机制增强特征选择能力。在LINEMOD、Occluded LINEMOD和YCB-VIDEO数据集上的实验表明,该方法性能具有竞争力,在多种物体及复杂场景下均表现出一致改进。

原文摘要 · Abstract (English)

This paper proposes PoseLecTr, a graph-based encoder-decoder framework that integrates a novel Legendre convolution with attention mechanisms for six-degree-of-freedom (6-DOF) object pose estimation from monocular RGB images. Conventional learning-based approaches predominantly rely on grid-structured convolutions, which can limit their ability to model higher-order and long-range dependencies among image features, especially in cluttered or occluded scenes. PoseLecTr addresses this limitation by constructing a graph representation from image features, where spatial relationships are explicitly modeled through graph connectivity. The proposed framework incorporates a Legendre convolution layer to improve numerical stability in graph convolution, together with spatial-attention and self-attention distillation to enhance feature selection. Experiments conducted on the LINEMOD, Occluded LINEMOD, and YCB-VIDEO datasets demonstrate that our method achieves competitive performance and shows consistent improvements across a wide range of objects and scene complexities.

姿态估计图神经网络注意力机制6D位姿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。