arXiv:2602.07082cs.CVcs.AI2026-02被引 2

让小模型在设备端也能精准理解多帧视频中的空间关系。

MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation

  • 通过构建全局语义地图整合多帧碎片化空间信息
  • 在资源受限设备上显著提升跨帧空间推理准确率
  • 适合部署在机器人等边缘计算场景的视觉导航任务

当具身智能从传统的目标检测与识别扩展到机器人操作与动作规划等高级任务时,基于视频输入的视觉空间推理成为感知物体间空间关系并指导设备行为的关键。然而,现有视觉语言模型(VLMs)因缺乏对三维空间信息的理解,尤其在涉及多帧复杂空间关系的任务中表现薄弱。本文提出一种新的推理时计算技术——MosaicThinker,用于增强设备端小型VLM在复杂跨帧推理任务中的空间推理能力。其核心思想是将多帧中的碎片化空间信息融合为统一的全局语义地图,并通过视觉提示引导VLM在此地图上进行空间推理。实验表明,该方法可在资源受限的具身智能设备上显著提升跨帧空间推理的准确性,适用于多种类型和复杂度的推理任务。

原文摘要 · Abstract (English)

When embodied AI is expanding from traditional object detection and recognition to more advanced tasks of robot manipulation and actuation planning, visual spatial reasoning from the video inputs is necessary to perceive the spatial relationships of objects and guide device actions. However, existing visual language models (VLMs) have very weak capabilities in spatial reasoning due to the lack of knowledge about 3D spatial information, especially when the reasoning task involve complex spatial relations across multiple video frames. In this paper, we present a new inference-time computing technique for on-device embodied AI, namely \emph{MosaicThinker}, which enhances the on-device small VLM's spatial reasoning capabilities on difficult cross-frame reasoning tasks. Our basic idea is to integrate fragmented spatial information from multiple frames into a unified space representation of global semantic map, and further guide the VLM's spatial reasoning over the semantic map via a visual prompt. Experiment results show that our technique can greatly enhance the accuracy of cross-frame spatial reasoning on resource-constrained embodied AI devices, over reasoning tasks with diverse types and complexities.

具身智能空间推理边缘计算视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。