通过动态采样提升3D物体姿态估计精度,尤其在遮挡和噪声下表现更优。
Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation
- 基于SO(3)等变卷积网络与正激励点采样策略,自适应选择关键采样点。
- 在三个数据集上超越现有方法,尤其在未知姿态、高遮挡场景中提升显著。
- 适合需要高鲁棒性姿态估计的工业检测与机器人抓取场景。
学习3D形状的神经隐式场是快速发展的领域,可在任意分辨率下表示形状。由于灵活性强,该技术已成功应用于形状重建、新视角图像合成,并近期拓展至物体姿态估计。神经隐式场能建立相机空间与物体规范空间之间的密集对应关系,包括相机空间中未观测区域,显著提升复杂场景(如高度遮挡、新形状)下的姿态估计性能。然而,由于缺乏直接观测信号,预测未观测相机空间区域的规范坐标仍具挑战,模型依赖泛化能力导致不确定性高。因此,在整个相机空间密集采样可能产生误差,阻碍学习并降低性能。为此,我们提出结合SO(3)等变卷积隐式网络与正激励点采样(PIPS)策略的方法。SO(3)等变卷积隐式网络在任意查询位置估计点级属性,性能优于多数现有基线。PIPS策略根据输入动态确定采样位置,提升网络精度与训练效率。该方法在三个姿态估计数据集上优于当前最优水平,尤其在未知姿态、高遮挡、新几何和严重噪声等挑战性场景中表现突出。
原文摘要 · Abstract (English)
Learning neural implicit fields of 3D shapes is a rapidly emerging field that enables shape representation at arbitrary resolutions. Due to the flexibility, neural implicit fields have succeeded in many research areas, including shape reconstruction, novel view image synthesis, and more recently, object pose estimation. Neural implicit fields enable learning dense correspondences between the camera space and the object's canonical space-including unobserved regions in camera space-significantly boosting object pose estimation performance in challenging scenarios like highly occluded objects and novel shapes. Despite progress, predicting canonical coordinates for unobserved camera-space regions remains challenging due to the lack of direct observational signals. This necessitates heavy reliance on the model's generalization ability, resulting in high uncertainty. Consequently, densely sampling points across the entire camera space may yield inaccurate estimations that hinder the learning process and compromise performance. To alleviate this problem, we propose a method combining an SO(3)-equivariant convolutional implicit network and a positive-incentive point sampling (PIPS) strategy. The SO(3)-equivariant convolutional implicit network estimates point-level attributes with SO(3)-equivariance at arbitrary query locations, demonstrating superior performance compared to most existing baselines. The PIPS strategy dynamically determines sampling locations based on the input, thereby boosting the network's accuracy and training efficiency. Our method outperforms the state-of-the-art on three pose estimation datasets. Notably, it demonstrates significant improvements in challenging scenarios, such as objects captured with unseen pose, high occlusion, novel geometry, and severe noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。