arXiv:2410.18195cs.CVcs.RO2024-10NeurIPS被引 14

让智能体在真实场景中找用户专属物品,区分相似对象。

Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments

  • 用视觉参考图+文本描述引导智能体定位特定物品
  • 新任务PIN需在多实例中准确识别目标,挑战现有导航方法
  • 适合研究个性化导航与真实场景智能体的学者

近年来,室内环境中的视觉导航研究兴趣显著增长,得益于Gibson和Matterport3D等照片级真实感模拟环境的大规模导航数据集的出现。然而,这些数据集支持的导航任务通常局限于采集时环境中存在的物体,且未考虑用户特定物品可能与同类物品混淆、在环境中多个位置出现的真实场景。为此,我们提出新任务:个性化实例级导航(PIN),即让具身智能体通过区分同类别多个实例,找到并抵达特定个人物品。该任务配套新数据集PInNED,由添加了额外3D物体的逼真场景构成。每轮任务中,目标物品通过一组背景为中性的视觉参考图像和人工标注的文本描述提供给智能体。通过全面评估与分析,我们揭示了PIN任务的挑战,以及当前面向物体驱动导航的方法在模块化与端到端智能体上的性能与局限性。

原文摘要 · Abstract (English)

In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the navigation tasks supported by these datasets are often restricted to the objects present in the environment at acquisition time. Also, they fail to account for the realistic scenario in which the target object is a user-specific instance that can be easily confused with similar objects and may be found in multiple locations within the environment. To address these limitations, we propose a new task denominated Personalized Instance-based Navigation (PIN), in which an embodied agent is tasked with locating and reaching a specific personal object by distinguishing it among multiple instances of the same category. The task is accompanied by PInNED, a dedicated new dataset composed of photo-realistic scenes augmented with additional 3D objects. In each episode, the target object is presented to the agent using two modalities: a set of visual reference images on a neutral background and manually annotated textual descriptions. Through comprehensive evaluations and analyses, we showcase the challenges of the PIN task as well as the performance and shortcomings of currently available methods designed for object-driven navigation, considering modular and end-to-end agents.

视觉导航个性化具身智能体真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。