提出细粒度评估框架,诊断视觉语言导航模型在不同指令类型上的表现
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
- 基于上下文无关语法构建指令分类体系,结合大模型半自动生成数据
- 发现模型在数量理解上停滞、方向概念存在选择性偏差等关键问题
- 适合研究语言引导导航系统或评估模型泛化能力的学者参考
本研究提出一种针对视觉语言导航(VLN)任务的新评估框架,旨在以更细粒度的方式诊断当前模型在各类指令下的表现。框架基于任务的上下文无关语法(CFG),作为问题分解基础和指令类别设计的核心前提。我们提出一种借助大语言模型(LLMs)的半自动方法来构建CFG,进而归纳并生成涵盖五类主要指令的数据:方向变化、地标识别、区域识别、垂直移动和数量理解。对多种模型的分析揭示了显著的性能差异与反复出现的问题,包括数量理解能力停滞、对方向概念存在严重选择性偏差等,这些发现为未来语言引导导航系统的发展提供了重要依据。
原文摘要 · Abstract (English)
This study presents a novel evaluation framework for the Vision-Language Navigation (VLN) task. It aims to diagnose current models for various instruction categories at a finer-grained level. The framework is structured around the context-free grammar (CFG) of the task. The CFG serves as the basis for the problem decomposition and the core premise of the instruction categories design. We propose a semi-automatic method for CFG construction with the help of Large-Language Models (LLMs). Then, we induct and generate data spanning five principal instruction categories (i.e. direction change, landmark recognition, region recognition, vertical movement, and numerical comprehension). Our analysis of different models reveals notable performance discrepancies and recurrent issues. The stagnation of numerical comprehension, heavy selective biases over directional concepts, and other interesting findings contribute to the development of future language-guided navigation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。