构建跨尺度空间推理数据集与模型,支持从毫米到公里的多场景理解。
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- 用自动化管道构建5个尺度的38000段视频场景,生成百万级问答对。
- 提出基于尺度感知的渐进式训练框架,有效缓解知识冲突问题。
- 适合机器人、自动驾驶等需要跨尺度空间理解的研究者使用。
随着空间推理研究的兴起,尽管在室内场景理解方面取得显著进展,但在机器人、自动驾驶等多样化应用中仍面临挑战。本文旨在通过解决两大关键问题来推进跨尺度空间推理:一是依赖室内3D扫描和人工标注的数据集构建;二是缺乏有效的全尺度场景建模,易导致对单个场景过拟合。为此,我们提出首个整合结构化空间推理知识体系、尺度感知建模与渐进式训练范式的整体方案,首次将多尺度视觉空间智能扩展至大语言模型。利用任务特异性的专家驱动自动化流程,我们构建了覆盖5个空间尺度的38,000段视频场景,形成包含约100万条空间问答对的SpaceVista-1M数据集,涵盖19种任务类型。虽然专家模型可注入领域知识,但其评估不可靠。因此,我们通过手动录制、检索与组装视频数据,建立精确标注的全尺度基准。然而,直接使用SpaceVista-1M训练效果不佳,存在知识冲突风险。为此,我们提出SpaceVista-7B模型,接受超越语义的密集输入,并以尺度为锚点引入尺度感知专家与渐进奖励机制。在5个基准测试(含自建SpaceVista-Bench)上的广泛评估表明,该模型表现优异,展现出强大的跨尺度与跨场景泛化能力。相关数据集、模型与基准将开源于https://peiwensun2000.github.io/mm2km。
原文摘要 · Abstract (English)
With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to advance all-scale spatial reasoning across diverse scenarios by tackling two key challenges: 1) the heavy reliance on indoor 3D scans and labor-intensive manual annotations for dataset curation; 2) the absence of effective all-scale scene modeling, which often leads to overfitting to individual scenes. In this paper, we introduce a holistic solution that integrates a structured spatial reasoning knowledge system, scale-aware modeling, and a progressive training paradigm, as the first attempt to broaden the all-scale spatial intelligence of MLLMs to the best of our knowledge. Using a task-specific, specialist-driven automated pipeline, we curate over 38K video scenes across 5 spatial scales to create SpaceVista-1M, a dataset comprising approximately 1M spatial QA pairs spanning 19 diverse task types. While specialist models can inject useful domain knowledge, they are not reliable for evaluation. We then build an all-scale benchmark with precise annotations by manually recording, retrieving, and assembling video-based data. However, naive training with SpaceVista-1M often yields suboptimal results due to the potential knowledge conflict. Accordingly, we introduce SpaceVista-7B, a spatial reasoning model that accepts dense inputs beyond semantics and uses scale as an anchor for scale-aware experts and progressive rewards. Finally, extensive evaluations across 5 benchmarks, including our SpaceVista-Bench, demonstrate competitive performance, showcasing strong generalization across all scales and scenarios. Our dataset, model, and benchmark will be released on https://peiwensun2000.github.io/mm2km .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。