arXiv:2603.17980cs.CV2026-03

用惯性数据提升视频3D理解,让模型更准更快

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding

  • 结合惯性数据与视觉信息筛选关键帧,减少计算量
  • 在多个3D任务中达到顶尖精度,速度比现有方法快1.61倍
  • 适合需要高效精准空间推理的自动驾驶等场景

近期多模态大语言模型在3D场景空间推理方面展现出巨大潜力,但通常依赖计算成本高昂的3D表示(如点云或重建的鸟瞰图),或缺乏物理尺度约束导致歧义。本文通过同步采集视频与惯性测量单元(IMU)数据,提出Motion-MLLM框架,包含两个核心组件:(1) 一种级联运动-视觉关键帧过滤模块,利用IMU数据和视觉特征高效选择稀疏但具有代表性的关键帧;(2) 一种非对称跨模态融合模块,以运动标记为中介,将运动线索与跨帧视觉上下文注入视觉表征。通过将视觉内容与真实运动轨迹对齐,Motion-MLLM可准确推断绝对尺度与空间关系。大量实验表明,该方法在多项3D场景理解与空间推理任务中显著优于现有方法,在保持竞争力精度的同时,分别实现1.30倍和1.61倍的速度提升。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye View (BEV) maps, or lack physical grounding to resolve ambiguities in scale and size. This paper significantly enhances MLLMs with egomotion modality data, captured by Inertial Measurement Units (IMUs) concurrently with the video. In particular, we propose a novel framework, called Motion-MLLM, introducing two key components: (1) a cascaded motion-visual keyframe filtering module that leverages both IMU data and visual features to efficiently select a sparse yet representative set of keyframes, and (2) an asymmetric cross-modal fusion module where motion tokens serve as intermediaries that channel egomotion cues and cross-frame visual context into the visual representation. By grounding visual content in physical egomotion trajectories, Motion-MLLM can reason about absolute scale and spatial relationships across the scene. Our extensive evaluation shows that Motion-MLLM makes significant improvements in various tasks related to 3D scene understanding and spatial reasoning. Compared to state-of-the-art (SOTA) methods based on video frames and explicit 3D data, Motion-MLLM achieves competitive accuracy while running $1.30\times$ and $1.61\times$ faster, respectively.

3D理解多模态IMU效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。