针对杂乱场景,提出基于物体感知的视角选择策略,提升3D重建精度。
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
- 利用物体特征生成置信度加权的信息增益,指导新视角选择。
- 在合成数据上深度误差降低77.14%,真实数据上降低34.10%。
- 聚焦目标物体的视角选择,使该物体误差再降25.60%,适合机器人抓取任务。
在存在遮挡和观测不完整的问题场景中,选择有信息量的视角对构建可靠三维表示至关重要。3D高斯点阵(3DGS)因其可显式引导后续视角选择并利用新观测优化表示而具备优势。然而,现有方法仅依赖几何线索,忽略与操作相关的语义信息,且倾向于过度利用而非探索。为此,本文提出一种实例感知的下一最佳视角(NBV)策略,通过利用物体特征优先探索未充分观测区域。具体地,所提物体感知3DGS将实例级信息提炼为独热编码的物体向量,用于计算置信度加权的信息增益,从而识别出存在错误或不确定性的高斯点区域。此外,该方法可轻松扩展为以物体为中心的NBV,聚焦于目标物体,提升对物体位置变化的重建鲁棒性。实验表明,与基线相比,本方法在合成数据集上深度误差降低77.14%,在真实世界GraspNet数据集上降低34.10%;相较于全局视角选择,针对特定物体执行NBV可使该物体深度误差进一步减少25.60%。我们还在真实机器人操作任务中验证了该方法的有效性。
原文摘要 · Abstract (English)
In cluttered scenes with inevitable occlusions and incomplete observations, selecting informative viewpoints is essential for building a reliable representation. In this context, 3D Gaussian Splatting (3DGS) offers a distinct advantage, as it can explicitly guide the selection of subsequent viewpoints and then refine the representation with new observations. However, existing approaches rely solely on geometric cues, neglect manipulation-relevant semantics, and tend to prioritize exploitation over exploration. To tackle these limitations, we introduce an instance-aware Next Best View (NBV) policy that prioritizes underexplored regions by leveraging object features. Specifically, our object-aware 3DGS distills instancelevel information into one-hot object vectors, which are used to compute confidence-weighted information gain that guides the identification of regions associated with erroneous and uncertain Gaussians. Furthermore, our method can be easily adapted to an object-centric NBV, which focuses view selection on a target object, thereby improving reconstruction robustness to object placement. Experiments demonstrate that our NBV policy reduces depth error by up to 77.14% on the synthetic dataset and 34.10% on the real-world GraspNet dataset compared to baselines. Moreover, compared to targeting the entire scene, performing NBV on a specific object yields an additional reduction of 25.60% in depth error for that object. We further validate the effectiveness of our approach through real-world robotic manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。