arXiv:2605.10106cs.CVcs.AI2026-05

无需训练,用视频引导大模型进行三维空间推理。

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models

论文配图:ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
图 1 · 摘自论文原文
  • 利用专家模型的显式空间信息构建可插拔推理框架
  • 在多个基准和未见任务上提升15.6%~28.9%准确率
  • 适合研究空间推理机制与降低数据标注成本的场景

多模态大语言模型(MLLMs)虽在3D空间智能方面取得进展,但主要依赖后训练于精心设计的基准,推理阶段方法仍较薄弱。本文提出无需训练的ViSRA——基于视频的空间推理代理,用于探测MLLMs的空间推理能力。ViSRA通过引入专家模型的显式空间信息,以模块化、可扩展方式激发空间推理,实现即插即用的灵活范式。其两大优势:(1) 获得人类对齐且可迁移的3D理解,避免任务特异性过拟合;(2) 无需后训练计算开销,也无须大量人工标注的空间推理数据集。实验表明,ViSRA在多个现有基准和未见3D空间推理任务上均带来一致提升,性能优于基线最高达15.6%和28.9%绝对差距。

原文摘要 · Abstract (English)

Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leaving the inference-time approach relatively underexplored. In this paper, we take a training-free perspective and introduce ViSRA, a human-aligned Video-based Spatial Reasoning Agent, as a framework to probe the spatial reasoning mechanism of MLLMs. ViSRA elicits spatial reasoning in a modular and extensible manner by leveraging explicit spatial information from expert models, enabling a plug-and-play flexible paradigm. ViSRA offers two key advantages: (1) human-aligned and transferable 3D understanding rather than task-specific overfitting; and (2) no post-training computational cost along with heavy manual curation of spatial reasoning datasets. Experimental results demonstrate consistent improvement across a set of MLLMs on both existing benchmarks and unseen 3D spatial reasoning tasks, with ViSRA outperforming baselines by up to a 15.6% and 28.9% absolute margin respectively.

空间推理多模态视频理解零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。