arXiv:2505.06219cs.CVcs.RO2025-05被引 4

直接优化重建质量,让机器人选视角更准

VIN-NBV: A View Introspection Network for Next-Best-View Selection

  • 用轻量神经网络预测新视角的重建提升效果
  • 相比传统方法,重建质量提升约30%
  • 适合无先验知识、有遮挡的复杂场景

Next Best View (NBV) 算法旨在以最少资源(如采集次数、时间或移动距离)最大化3D场景重建质量。以往方法常以覆盖率作为重建质量的代理指标,但在存在遮挡和细节复杂的场景中,该策略往往不足,导致重建效果差。本文提出关键洞察:应直接训练采集策略以优化重建质量,而非仅追求覆盖。为此,我们引入视图内省网络(View Introspection Network, VIN),一个轻量级神经网络,能在不进行任何新采集的情况下预测潜在下一视角的相对重建提升(RRI)。基于此,设计了一种简单但高效的基于序列采样的贪心式NBV策略。所提方法VIN-NBV具有泛化能力,可适应未见物体类别,无需场景先验知识,且能灵活应对资源约束与遮挡问题。实验表明,采用RRI评估标准,在相同贪心策略下,重建质量相较覆盖率标准提升约30%;同时,优于深度强化学习方法Scan-RL和GenNBV约40%。

原文摘要 · Abstract (English)

Next Best View (NBV) algorithms aim to maximize 3D scene acquisition quality using minimal resources, e.g. number of acquisitions, time taken, or distance traversed. Prior methods often rely on coverage maximization as a proxy for reconstruction quality, but for complex scenes with occlusions and finer details, this is not always sufficient and leads to poor reconstructions. Our key insight is to train an acquisition policy that directly optimizes for reconstruction quality rather than just coverage. To achieve this, we introduce the View Introspection Network (VIN): a lightweight neural network that predicts the Relative Reconstruction Improvement (RRI) of a potential next viewpoint without making any new acquisitions. We use this network to power a simple, yet effective, sequential samplingbased greedy NBV policy. Our approach, VIN-NBV, generalizes to unseen object categories, operates without prior scene knowledge, is adaptable to resource constraints, and can handle occlusions. We show that our RRI fitness criterion leads to a ~30% gain in reconstruction quality over a coverage-based criterion using the same greedy strategy. Furthermore, VIN-NBV also outperforms deep reinforcement learning methods, Scan-RL and GenNBV, by ~40%.

3D重建视角选择神经网络智能采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。