用大模型实现多样场景下的开放词汇物体导航,效果超越GPT-4o
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
- 构建大规模数据集DivScene,覆盖81类场景与5707种目标物体
- 仅用BFS生成路径微调,导航成功率超GPT-4o 20%以上
- 无需人工标注,基于思维链推理提升模型导航能力
大型视觉语言模型(LVLM)在视觉问答和文档理解任务中取得显著进展,但其在具身环境中的理解与导航潜力尚未充分探索。本文提出DivScene,一个包含4,614栋房屋、81种场景类型和5,707种目标物体的大规模数据集,相较现有数据集具有更高的目标与场景多样性,支持全面评估开放词汇物体导航任务。我们在该数据集上评估了多种基于LVLM和LLM的方法,发现当前模型仍难以胜任开放词汇导航。随后,我们对LVLM进行微调,使其基于思维链(CoT)生成下一步动作。结果显示,仅使用BFS生成的最短路径进行无监督微调,即可显著提升导航性能,成功率达GPT-4o的20%以上。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains underexplored. In this work, we first study the challenge of open-vocabulary object navigation by introducing DivScene, a large-scale dataset with 4,614 houses across 81 scene types and 5,707 kinds of target objects. Our dataset provides a much greater diversity of target objects and scene types than existing datasets, enabling a comprehensive task evaluation. We evaluated various methods with LVLMs and LLMs on our dataset and found that current models still fall short of open-vocab object navigation ability. Then, we fine-tuned LVLMs to predict the next action with CoT explanations. We observe that LVLM's navigation ability can be improved substantially with only BFS-generated shortest paths without any human supervision, surpassing GPT-4o by over 20% in success rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。