arXiv:2605.27367cs.CV2026-05

评测空间大模型跨任务、跨视角的通用能力,发现当前模型仍非全能选手。

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

论文配图:SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
图 1 · 摘自论文原文
  • 构建跨范式、多场景的基准 SpatialBench,覆盖546个场景和19个数据集
  • 实验证明全上下文注意力精度最高,有限内存策略更利于长序列扩展
  • 强调真实场景对齐与数据质量比单纯数据量更重要,适合模型泛化研究者

尽管空间基础模型在标准数据集上表现优异,但一个关键问题仍未解决:它们是否真正具备跨多样化下游任务、任意视角、场景域变化、输入密度差异及硬件约束的鲁棒泛化能力?现有评估大多局限于特定领域,受限于窄范式覆盖、有限场景域和任意帧采样,难以真实衡量其泛化性能。为此,我们提出 SpatialBench,一个跨范式、域多样且采用确定性采样的基准。它包含19个数据集和546个场景,覆盖5类空间域,全面评估41个模型在6种范式下的5个任务套件,涵盖4种输入密度设置。评估表明当前模型尚非全能选手,并揭示关键洞见:全上下文注意力最大化精度,而有界记忆策略提升长序列可扩展性。此外,在具身与第一人称任务中,严格域对齐和高质量数据比单纯数据量更关键。为填补最大数据缺口,我们引入大规模数据集 DA-Next-5M 和强基线模型 DA-Next,推动空间表征学习边界。

原文摘要 · Abstract (English)

While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round players capable of generalizing robustly across diverse downstream tasks, arbitrary viewpoints, shifting scene domains, varying input densities, and specific hardware constraints? Answering this overarching question requires a holistic assessment, yet current models are mainly evaluated on specific domains for which they were specifically designed or trained. Such evaluations are intrinsically limited by narrow paradigm coverage, limited scene domains, and arbitrary frame sampling, making it fundamentally difficult to assess their true generalization capabilities. To address this gap, we present SpatialBench, a cross-paradigm, domain-diverse benchmark for spatial foundation models with deterministic sampling. SpatialBench features unprecedented scale and rigorous deterministic design, comprising 19 datasets and 546 scenes across 5 diverse spatial domains. It comprehensively evaluates 41 models across 6 paradigms on 5 task suites under 4 different input density settings. Our extensive evaluation reveals that current models are not yet all-round players, and uncovers crucial insights for future advancement. Specifically, we demonstrate that full-context attention maximizes accuracy while bounded-memory strategies unlock long-sequence scalability. Moreover, our empirical evaluations in challenging embodied and egocentric tasks demonstrate that strict domain alignment and high data quality are far more critical to performance than simple dataset scaling. Furthermore, to address the largest data gap identified in our analysis, we go beyond evaluation by introducing a large-scale dataset, DA-Next-5M, and a strong baseline model, DA-Next, pushing the boundaries of spatial representation learning.

空间模型基准评测泛化能力数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。