arXiv:2504.08603cs.ROcs.AI2025-04被引 14

让机器人在任意环境实时理解物体并建图,支持开放词汇查询。

FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment

  • 用视觉语言模型提取物体级特征,融合到3D体素子地图中
  • 相比现有方法,速度更快、内存占用更低,支持大规模部署
  • 适合无人机等资源受限设备,可用于搜救等实际任务

几何准确且语义丰富的地图表示对机器人在未知环境中的部署与任务规划至关重要。然而,实时地对大规模未知环境进行开放词汇的语义理解仍面临挑战,主要源于计算开销。本文提出 FindAnything,一种将视觉-语言信息融入密集体素子地图的开放世界建图框架。通过视觉-语言特征,该方法结合纯几何与开放词汇语义信息,实现更高层次的理解。其创新在于以物体为中心聚合特征:基于eSAM分割结果,将像素级视觉-语言特征聚合成对象级表示,并集成至体素子地图,实现从开放词汇查询到3D几何的映射,且内存效率高。实验表明,FindAnything 在语义精度上达到当前最优水平,同时显著提升速度与内存效率,可在大规模环境及资源受限设备(如微型无人机)上部署。我们进一步验证其实时能力适用于下游任务,如模拟搜救场景中的自主微型无人机探索。

原文摘要 · Abstract (English)

Geometrically accurate and semantically expressive map representations have proven invaluable for robot deployment and task planning in unknown environments. Nevertheless, real-time, open-vocabulary semantic understanding of large-scale unknown environments still presents open challenges, mainly due to computational requirements. In this paper we present FindAnything, an open-world mapping framework that incorporates vision-language information into dense volumetric submaps. Thanks to the use of vision-language features, FindAnything combines pure geometric and open-vocabulary semantic information for a higher level of understanding. It proposes an efficient storage of open-vocabulary information through the aggregation of features at the object level. Pixelwise vision-language features are aggregated based on eSAM segments, which are in turn integrated into object-centric volumetric submaps, providing a mapping from open-vocabulary queries to 3D geometry that is scalable also in terms of memory usage. We demonstrate that FindAnything performs on par with the state-of-the-art in terms of semantic accuracy while being substantially faster and more memory-efficient, allowing its deployment in large-scale environments and on resourceconstrained devices, such as MAVs. We show that the real-time capabilities of FindAnything make it useful for downstream tasks, such as autonomous MAV exploration in a simulated Search and Rescue scenario. Project Page: https://ethz-mrl.github.io/findanything/.

机器人语义建图开放词汇无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。