arXiv:2506.02896cs.CVcs.LG2025-06NeurIPS被引 6

测试视觉语言模型在复杂环境中的探索能力,发现其表现远不如人类。

FlySearch: Exploring how vision-language models explore

  • 构建3D户外真实场景搜索环境FlySearch,评估模型主动探索能力。
  • 顶尖视觉语言模型在简单任务上成功率不足50%,与人类差距随难度增大而拉大。
  • 发现模型幻觉、理解偏差和规划失败是主要问题,微调可部分改善。

现实世界杂乱无序,获取关键信息常需主动、目标驱动的探索。当前主流的视觉语言模型(VLM)虽在多项零样本任务中表现出色,但其在动态复杂环境下的探索能力尚不明确。本文提出FlySearch——一个3D、室外、照片级真实的搜寻与导航环境,用于在复杂场景中定位物体。设计三类难度递增的场景,结果显示,最先进的VLM在最简单任务上也难以稳定完成,且与人类表现的差距随任务复杂度上升而扩大。我们识别出核心原因:从视觉幻觉、上下文误解到任务规划失败。实验表明,部分问题可通过微调缓解。相关基准、场景与代码已开源。

原文摘要 · Abstract (English)

The real world is messy and unstructured. Uncovering critical information often requires active, goal-driven exploration. It remains to be seen whether Vision-Language Models (VLMs), which recently emerged as a popular zero-shot tool in many difficult tasks, can operate effectively in such conditions. In this paper, we answer this question by introducing FlySearch, a 3D, outdoor, photorealistic environment for searching and navigating to objects in complex scenes. We define three sets of scenarios with varying difficulty and observe that state-of-the-art VLMs cannot reliably solve even the simplest exploration tasks, with the gap to human performance increasing as the tasks get harder. We identify a set of central causes, ranging from vision hallucination, through context misunderstanding, to task planning failures, and we show that some of them can be addressed by finetuning. We publicly release the benchmark, scenarios, and the underlying codebase.

视觉语言模型主动探索3D导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。