arXiv:2510.06754cs.ROcs.CV2025-10

统一建模视觉语义与空间不确定性,让机器人在新环境自主探索时更可靠。

UniFField: A Generalizable Unified Neural Feature Field for Visual, Semantic, and Spatial Uncertainties in Any Scene

  • 用体素网格融合多模态特征,实现零样本泛化。
  • 实时更新场景特征与预测不确定性,误差评估准确率高。
  • 适合需可靠决策的机器人导航与操作任务。

三维场景的视觉、几何与语义理解对机器人在非结构化复杂环境中的任务执行至关重要。为实现鲁棒决策,机器人需评估感知信息的可靠性。尽管近期3D神经特征场已使机器人利用预训练基础模型完成语言引导操作与导航,但现有方法存在两大局限:(i) 通常仅适用于特定场景,(ii) 缺乏对预测不确定性的建模能力。本文提出UniFField,一种统一的、具备不确定性感知的神经特征场,将视觉、语义与几何特征整合于单一可泛化的表示中,并同时预测各模态的不确定性。该方法可零样本应用于任意新环境,随着机器人探索过程,逐步将RGB-D图像融入体素化特征表示,同步更新不确定性估计。我们验证了不确定性估计能准确描述场景重建与语义特征预测中的模型误差。此外,成功利用特征预测及其不确定性,在移动操作机器人上实现了主动目标搜索任务,展现了其在鲁棒决策中的能力。

原文摘要 · Abstract (English)

Comprehensive visual, geometric, and semantic understanding of a 3D scene is crucial for successful execution of robotic tasks, especially in unstructured and complex environments. Additionally, to make robust decisions, it is necessary for the robot to evaluate the reliability of perceived information. While recent advances in 3D neural feature fields have enabled robots to leverage features from pretrained foundation models for tasks such as language-guided manipulation and navigation, existing methods suffer from two critical limitations: (i) they are typically scene-specific, and (ii) they lack the ability to model uncertainty in their predictions. We present UniFField, a unified uncertainty-aware neural feature field that combines visual, semantic, and geometric features in a single generalizable representation while also predicting uncertainty in each modality. Our approach, which can be applied zero shot to any new environment, incrementally integrates RGB-D images into our voxel-based feature representation as the robot explores the scene, simultaneously updating uncertainty estimation. We evaluate our uncertainty estimations to accurately describe the model prediction errors in scene reconstruction and semantic feature prediction. Furthermore, we successfully leverage our feature predictions and their respective uncertainty for an active object search task using a mobile manipulator robot, demonstrating the capability for robust decision-making.

3D理解不确定性建模机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。