让大模型用参照物推理空间关系,性能提升超40点。
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
- 通过引导模型使用参照物构建推理路径,提升空间量化能力。
- 最佳模型在引入参照物后准确率提高19点,其他模型提升超30点。
- 无需数据、架构或微调,零样本提示即可显著增强推理效果。
尽管视觉语言模型(VLMs)在描述图像中复杂关系方面取得进展,但其对物体尺寸和距离的定量推理能力仍待探索。本文构建了人工标注的基准Q-Spatial Bench,包含271个问题,涵盖五个类别,用于系统评估当前先进VLMs的定量空间推理能力。分析显示,对象间距离推理尤为困难;部分VLM表现显著优于其他模型,最佳与次佳模型间准确率差距超过40点。令人意外的是,当顶级模型自然生成以参照物为依据的推理路径时,成功率提升19点。基于此,我们提出零样本提示方法SpatialPrompt,指导模型在回答中使用参照物。应用该方法后,Gemini 1.5 Pro、Gemini 1.5 Flash和GPT-4V准确率分别提升40、20、30点以上,且无需额外数据、模型修改或微调。
原文摘要 · Abstract (English)
Despite recent advances demonstrating vision-language models' (VLMs) abilities to describe complex relationships in images using natural language, their capability to quantitatively reason about object sizes and distances remains underexplored. In this work, we introduce a manually annotated benchmark, Q-Spatial Bench, with 271 questions across five categories designed for quantitative spatial reasoning and systematically investigate the performance of state-of-the-art VLMs on this task. Our analysis reveals that reasoning about distances between objects is particularly challenging for SoTA VLMs; however, some VLMs significantly outperform others, with an over 40-point gap between the two best performing models. We also make the surprising observation that the success rate of the top-performing VLM increases by 19 points when a reasoning path using a reference object emerges naturally in the response. Inspired by this observation, we develop a zero-shot prompting technique, SpatialPrompt, that encourages VLMs to answer quantitative spatial questions using reference objects as visual cues. By instructing VLMs to use reference objects in their reasoning paths via SpatialPrompt, Gemini 1.5 Pro, Gemini 1.5 Flash, and GPT-4V improve their success rates by over 40, 20, and 30 points, respectively. We emphasize that these significant improvements are obtained without needing more data, model architectural modifications, or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。