让多模态大模型通过文本推理空间关系,提升视频理解准确率
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
- 用文本描述3D环境结构,引导模型分步推理空间关系
- 在两个基准测试上显著超越现有方法,跨不同模型均有效
- 适合研究视频理解与空间推理的学者和开发者
现有多模态大语言模型在3D空间推理方面表现不佳,因无法构建视频中描绘的3D环境的结构化抽象。受外源性空间推理认知理论启发,本文提出一种名为TRACE的提示方法,引导模型生成视频的文本化空间表示作为中间推理过程,以更准确回答空间问题。TRACE编码元上下文、相机轨迹及具体物体实体,支持对第一人称视频的结构化空间推理。在VSI-Bench和OST-Bench上的大量实验表明,TRACE在多种不同参数规模和训练范式的模型上均取得显著且一致的提升。进一步的消融实验验证了设计选择的有效性,并深入分析了模型在3D空间推理中的瓶颈。
原文摘要 · Abstract (English)
Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inputs. To bridge this gap, drawing inspiration from cognitive theories of allocentric spatial reasoning, we investigate how to enable MLLMs to model and reason over text-based spatial representations of video. Specifically, we introduce Textual Representation of Allocentric Context from Egocentric Video (TRACE), a prompting method that induces MLLMs to generate text-based representations of 3D environments as intermediate reasoning traces for more accurate spatial question answering. TRACE encodes meta-context, camera trajectories, and detailed object entities to support structured spatial reasoning over egocentric videos. Extensive experiments on VSI-Bench and OST-Bench demonstrate that TRACE yields notable and consistent improvements over prior prompting strategies across a diverse range of MLLM backbones, spanning different parameter scales and training schemas. We further present ablation studies to validate our design choices, along with detailed analyses that probe the bottlenecks of 3D spatial reasoning in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。