无需训练,通过视觉多样性与运动重建提升多模态模型空间理解能力。
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- 采用语义丰富采样和视觉轨迹模拟,增强图像空间信息。
- 在多个空间推理任务上提升模型性能,最高增益达3.5%。
- 无需训练和额外计算,可直接接入现有多模态大模型。
我们提出SEE&TREK,首个面向纯视觉约束下提升多模态大语言模型(MLLMs)空间理解能力的免训练提示框架。现有方法多依赖深度或点云等模态,而纯视觉空间理解仍被忽视。该框架基于两大核心原则:提升视觉多样性与重建运动轨迹。为增强视觉多样性,采用最大语义丰富度采样,利用现成感知模型提取蕴含场景结构的语义关键帧;为实现运动重建,模拟视觉轨迹并将相对空间位置编码至关键帧,以保持空间关系与时间连贯性。方法完全免训练且无需GPU,仅需一次前向传播,可无缝集成至现有MLLM。在VSI-BENCH与STI-BENCH上的大量实验表明,其在多种空间推理任务中持续提升性能,最高提升达+3.5%,为强化空间智能提供了有效路径。
原文摘要 · Abstract (English)
We introduce SEE&TREK, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMS) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visualspatial understanding remains underexplored. SEE&TREK addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPU-free, requiring only a single forward pass, and can be seamlessly integrated into existing MLLM'S. Extensive experiments on the VSI-B ENCH and STI-B ENCH show that S EE &T REK consistently boosts various MLLM S performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。