arXiv:2605.22536cs.CVcs.CL2026-05

首个针对视觉退化场景的空间智能评测数据集,揭示大模型在真实环境下的脆弱性。

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

论文配图:SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
图 1 · 摘自论文原文
  • 用物理建模的退化生成引擎,在3D高斯点云渲染中模拟9类真实退化
  • 构建100万条问答对,覆盖11类空间推理任务与9种退化类型
  • 微调后模型在退化环境下超越人类表现,且不影响清晰图像性能

多模态大语言模型在空间智能方面进展迅速,但现有评测基准大多假设视觉输入完好,忽视了实际部署中常见的退化问题,如运动模糊、低光、恶劣天气、镜头畸变和压缩伪影。这引发了一个根本问题:当前多模态大语言模型在视觉观测不完美时,其空间智能有多稳健?为此,我们提出了SpaceDG,首个面向退化感知的空间理解大规模数据集。该数据集基于物理驱动的退化合成引擎,将退化形成过程嵌入3D高斯点阵(3DGS)渲染,可真实模拟九类退化。最终数据集包含约100万条来自近1000个室内场景的问答对。我们进一步构建了SpaceDG-Bench,一个经人工验证的评测基准,涵盖11类推理任务与9种视觉退化,共产生超过10,000个视觉问答实例。对25个开源与闭源多模态大模型的评估显示,视觉退化会持续且显著削弱空间推理能力,暴露出严重的鲁棒性缺口。最后,我们证明在SpaceDG上微调能显著提升退化鲁棒性,甚至在退化条件下超越人类表现,且在干净图像上无性能下降,表明退化感知训练对构建鲁棒空间智能具有巨大潜力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pristine visual inputs and overlook the degradations that commonly occur in real-world deployment, such as motion blur, low light, adverse weather, lens distortion, and compression artifacts. This raises a fundamental question: how robust is the spatial intelligence of current MLLMs when visual observations are imperfect? To answer this question, we introduce SpaceDG, the first large-scale dataset for degradation-aware spatial understanding. It is constructed with a physically grounded degradation synthesis engine that embeds degradation formation process into 3D Gaussian Splatting (3DGS) rendering, enabling realistic simulation of nine degradation types. The resulting dataset contains approximately 1M QA pairs from nearly 1,000 indoor scenes. We further introduce SpaceDG-Bench, an human-verified benchmark with 1,102 questions spanning 11 reasoning categories and 9 visual degradation types, yielding over 10K VQA instances. Evaluating 25 open- and closed-source MLLMs reveals that visual degradations consistently and substantially impair spatial reasoning, exposing a critical robustness gap. Finally, we show that finetuning on SpaceDG markedly improves degradation robustness and can even surpass human performance under degraded conditions without any performance drop on clean images, highlighting the promise of degradation-aware training for robust spatial intelligence.

空间推理视觉退化大模型评测3DGS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。