用全局运动引导流场,解决遮挡下的6D位姿估计难题
GMFlow: Global Motion-Guided Recurrent Flow for 6D Object Pose Estimation
- 通过线性注意力捕捉全局上下文,引导局部运动推断整体位姿
- 在LM-O和YCB-V数据集上精度优于现有方法,计算效率不降低
- 适合处理部分遮挡或缺失的刚体物体位姿估计任务
6D物体位姿估计对机器人感知与精准操作至关重要。遮挡和物体可见性不完整是常见挑战,但现有位姿精修方法往往难以有效应对。为此,我们提出一种全局运动引导的递归光流估计方法GMFlow,用于位姿估计。GMFlow通过寻求全局解释来克服遮挡或缺失部分带来的局部歧义,利用物体结构信息将可见部分的运动拓展至不可见区域。具体地,通过线性注意力机制捕捉全局上下文信息,并引导局部运动生成全局运动估计。此外,在光流迭代过程中引入物体形状约束,使光流估计更适用于位姿估计场景。在LM-O和YCB-V数据集上的实验表明,该方法在保持竞争力的计算效率的同时,显著提升了精度。
原文摘要 · Abstract (English)
6D object pose estimation is crucial for robotic perception and precise manipulation. Occlusion and incomplete object visibility are common challenges in this task, but existing pose refinement methods often struggle to handle these issues effectively. To tackle this problem, we propose a global motion-guided recurrent flow estimation method called GMFlow for pose estimation. GMFlow overcomes local ambiguities caused by occlusion or missing parts by seeking global explanations. We leverage the object's structural information to extend the motion of visible parts of the rigid body to its invisible regions. Specifically, we capture global contextual information through a linear attention mechanism and guide local motion information to generate global motion estimates. Furthermore, we introduce object shape constraints in the flow iteration process, making flow estimation suitable for pose estimation scenarios. Experiments on the LM-O and YCB-V datasets demonstrate that our method outperforms existing techniques in accuracy while maintaining competitive computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。