arXiv:2511.04668cs.CV2025-11被引 16

用3D模拟生成空间视频数据,提升模型跨时空推理能力

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

  • 借助3D模拟器生成带精确空间标注的视频数据
  • 仅25K条模拟数据即超越72B大模型表现
  • 适合需高效训练空间理解能力的研究者

尽管多模态语言模型在高层视频理解上表现优异,但在时空交叉的空间推理方面仍存在不足。当前的空间训练依赖真实视频数据,但获取多样化且带有精确空间标注的视频仍是瓶颈。为此,我们提出SIMS-V——一种利用3D模拟器优势生成丰富空间信息视频数据的系统化框架。通过系统消融实验,我们研究了问题类型、混合方式与规模对真实世界迁移效果的影响,发现仅需三种问题类别(度量测量、视角相关推理、时间追踪)即可有效培养可迁移的空间智能,尽管种类更少,但优于全面覆盖。这些发现使训练效率大幅提升:使用25,000个模拟样本微调的7B参数视频语言模型,性能超过72B基准模型,并在严格的真实世界空间推理基准上达到与专有模型相当的水平。该方法展现出强泛化能力,在保持通用视频理解性能的同时,显著提升具身与真实场景下的空间任务表现。

原文摘要 · Abstract (English)

Despite impressive high-level video comprehension, multimodal language models struggle with spatial reasoning across time and space. While current spatial training approaches rely on real-world video data, obtaining diverse footage with precise spatial annotations remains a bottleneck. To alleviate this bottleneck, we present SIMS-V -- a systematic data-generation framework that leverages the privileged information of 3D simulators to create spatially-rich video training data for multimodal language models. Using this framework, we investigate which properties of simulated data drive effective real-world transfer through systematic ablations of question types, mixes, and scales. We identify a minimal set of three question categories (metric measurement, perspective-dependent reasoning, and temporal tracking) that prove most effective for developing transferable spatial intelligence, outperforming comprehensive coverage despite using fewer question types. These insights enable highly efficient training: our 7B-parameter video LLM fine-tuned on just 25K simulated examples outperforms the larger 72B baseline and achieves competitive performance with proprietary models on rigorous real-world spatial reasoning benchmarks. Our approach demonstrates robust generalization, maintaining performance on general video understanding while showing substantial improvements on embodied and real-world spatial tasks.

空间推理3D模拟视频理解高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。