构建动态场景下开放词汇导航的仿真数据集与评估流水线
SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes
- 用多模态大模型生成符合现实语义的逼真场景变体
- 提供3000和10000个导航任务样本,支持可操作物体与动态环境
- 适合研究机器人导航、具身智能的学者与开发者使用
我们提出面向动态场景中开放词汇物体导航的语义感知数据集与基准生成流水线SD-OVON。该流水线利用预训练多模态基础模型生成无限数量的独特逼真场景变体,符合真实世界语义与日常常识,用于导航智能体的训练与评估,并配备适配Habitat模拟器的插件以生成导航任务实例。此外,我们提供了两个预先生成的任务数据集:SD-OVON-3k(约3000个任务)和SD-OVON-10k(约10000个任务),均源自包含2500个真实环境扫描的SD-OVON-Scenes数据集以及包含900个手动校验的扫描与艺术创作可操作物体模型的SD-OVON-Objects数据集。与以往局限于静态环境的数据集不同,SD-OVON涵盖动态场景与可操作物体,支持真实到仿真及仿真到真实的机器人应用。该方法提升了复杂环境下导航任务的真实性,强化了开放词汇导航智能体的训练与评估效果。为验证其有效性,我们提出了两个基线,并在SD-OVON-3k上与最先进方法进行对比。数据集、基准与源代码均已公开。
原文摘要 · Abstract (English)
We present the Semantics-aware Dataset and Benchmark Generation Pipeline for Open-vocabulary Object Navigation in Dynamic Scenes (SD-OVON). It utilizes pretraining multimodal foundation models to generate infinite unique photo-realistic scene variants that adhere to real-world semantics and daily commonsense for the training and the evaluation of navigation agents, accompanied with a plugin for generating object navigation task episodes compatible to the Habitat simulator. In addition, we offer two pre-generated object navigation task datasets, SD-OVON-3k and SD-OVON-10k, comprising respectively about 3k and 10k episodes of the open-vocabulary object navigation task, derived from the SD-OVON-Scenes dataset with 2.5k photo-realistic scans of real-world environments and the SD-OVON-Objects dataset with 0.9k manually inspected scanned and artist-created manipulatable object models. Unlike prior datasets limited to static environments, SD-OVON covers dynamic scenes and manipulatable objects, facilitating both real-to-sim and sim-to-real robotic applications. This approach enhances the realism of navigation tasks, the training and the evaluation of open-vocabulary object navigation agents in complex settings. To demonstrate the effectiveness of our pipeline and datasets, we propose two baselines and evaluate them along with state-of-the-art baselines on SD-OVON-3k. The datasets, benchmark and source code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。