让视觉语言模型学会在时空中思考,提升对3D空间和长视频的理解能力。
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- 通过构建三维空间与时间维度的视觉思维轨迹,让模型在回答前进行时空推理。
- 零样本下使GPT-4o在EgoSchema上提升15.8%,微调后在VSI-Bench上比Qwen2.5-VL-7B高8.3%。
- 仅用2D图像输入即可提升通用模型的时空理解能力,适合多模态研究者使用。
人类能从序列视觉观察中感知和理解三维空间及长视频。但视觉语言模型(VLMs)是否具备此能力?现有研究表明,即使最先进的VLMs在3D空间与长视频理解方面仍存在明显短板。当前方法通常需为3D任务和视频理解任务分别设计专用架构。本文提出LAST(LeArn to Think in Space and Time),仅需2D图像输入,即可联合提升通用VLMs对3D空间与长视频的理解能力。LAST使VLMs在给出最终答案前“思考”于空间与时间维度,构建三维空间与时间轴上的视觉思维轨迹。我们在两种场景验证其有效性:1)零样本提示专有模型;2)用包含时空思维轨迹的数据微调通用VLMs。结果表明,LAST在多个基准测试中均带来显著提升,涵盖3项空间理解、4项视频理解及3项图像理解任务。尤其在零样本下,使GPT-4o在EgoSchema上提升15.8%;在微调设置下,相比Qwen2.5-VL-7B在VSI-Bench上提高8.3%。
原文摘要 · Abstract (English)
Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-of-the-art VLMs still struggle to understand 3D space and long videos, although they are powerful in typical vision-language tasks. Current methods often rely on specialized architectural designs to improve performance for 3D tasks and video understanding tasks separately. In contrast, we propose LAST, short for LeArn to Think in Space and Time, to jointly improve 3D spatial and long video understanding for general VLMs with only a set of 2D images as inputs. LAST makes VLMs think in space and time rather than only with text before giving the final answer, building visual thinking trajectories in 3D space and temporal dimension. We demonstrate the effectiveness of LAST in two scenarios: 1) zero-shot, where we directly prompt proprietary models; and 2) fine-tuning general VLMs with data that include thinking trajectories in 3D space and time. We show that LAST brings substantial gains in various benchmarks, including 3 spatial understanding, 4 video understanding, and 3 image understanding tasks. Notably, 15.8% gains on EgoSchema with GPT-4o in a zero-shot manner and 8.3 gains on VSI-Bench compared with Qwen2.5-VL-7B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。