arXiv:2512.03040cs.CVcs.AI2025-12被引 2

仅用视频数据训练,模型就能完成导航与物体定位等空间推理任务。

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation

  • 基于视频上下文生成,无需深度或姿态等额外信息。
  • 能端到端规划路径并精准定位目标物体,保持空间一致性。
  • 适用于长序列和陌生环境,展现类人空间智能。

我们研究视频生成模型是否能具备人类认知核心的视觉空间智能,仅依赖视觉数据。为此,提出 Video4Spatial 框架,证明仅以视频场景为条件的视频扩散模型可完成复杂空间任务。验证了两个任务:场景导航——遵循相机位姿指令的同时保持场景三维几何一致性;物体定位——需语义定位、指令跟随与规划。两者均使用纯视频输入,无深度、位姿等辅助模态。通过简洁有效的框架设计与数据构建,Video4Spatial 展现出强大的视频上下文空间理解能力:实现端到端导航规划与目标定位,遵循相机指令并维持空间一致性,且在长上下文与域外环境中具有泛化能力。这些结果推动视频生成模型向通用视觉空间推理迈进。

原文摘要 · Abstract (English)

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models conditioned solely on video-based scene context can perform complex spatial tasks. We validate on two tasks: scene navigation - following camera-pose instructions while remaining consistent with 3D geometry of the scene, and object grounding - which requires semantic localization, instruction following, and planning. Both tasks use video-only inputs, without auxiliary modalities such as depth or poses. With simple yet effective design choices in the framework and data curation, Video4Spatial demonstrates strong spatial understanding from video context: it plans navigation and grounds target objects end-to-end, follows camera-pose instructions while maintaining spatial consistency, and generalizes to long contexts and out-of-domain environments. Taken together, these results advance video generative models toward general visuospatial reasoning.

视频生成空间推理扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。