重建3D空间智能评估基准,让视觉语言模型的三维推理更准确可信。
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

- 重新标注381个场景的物体与几何信息,修复原始数据缺陷
- 生成经人工验证的高质量问答对,确保模型输入下可回答
- 支持多帧预算和细粒度可见性控制,适合诊断模型缺陷
当前视觉语言模型(VLM)在三维空间智能评估中存在系统性无效问题。首先,许多基准使用基于点云的3D标注生成问答对,但这些标注在视频评估中可能遗漏可见物体、误标物体身份或扭曲依赖几何的答案(如尺寸),导致错误或模糊的问题。其次,评估常假设全场景可见,而多数VLM仅处理稀疏采样的帧(如16-64帧),使许多问题在实际输入下无法回答。为此,本文提出ReVSI,通过在5个数据集共381个场景中重新标注物体与几何信息,提升数据质量,并利用专业3D标注工具和人工验证重生成所有问答对,消除偏见。同时提供16/32/64/全部帧等多帧预算版本及细粒度物体可见性元数据,实现可控诊断分析。在ReVSI上对通用与领域特定VLM的评估揭示了以往基准掩盖的系统性失败模式,提供了更可靠、更具诊断性的空间智能评估。
原文摘要 · Abstract (English)
Current evaluations of spatial intelligence can be systematically invalid under modern vision-language model (VLM) settings. First, many benchmarks derive question-answer (QA) pairs from point-cloud-based 3D annotations originally curated for traditional 3D perception. When such annotations are treated as ground truth for video-based evaluation, reconstruction and annotation artifacts can miss objects that are clearly visible in the video, mislabel object identities, or corrupt geometry-dependent answers (e.g., size), yielding incorrect or ambiguous QA pairs. Second, evaluations often assume full-scene access, while many VLMs operate on sparsely sampled frames (e.g., 16-64), making many questions effectively unanswerable under the actual model inputs. We improve evaluation validity by introducing ReVSI, a benchmark and protocol that ensures each QA pair is answerable and correct under the model's actual inputs. To this end, we re-annotate objects and geometry across 381 scenes from 5 datasets to improve data quality, and regenerate all QA pairs with rigorous bias mitigation and human verification using professional 3D annotation tools. We further enhance evaluation controllability by providing variants across multiple frame budgets (16/32/64/all) and fine-grained object visibility metadata, enabling controlled diagnostic analyses. Evaluations of general and domain-specific VLMs on ReVSI reveal systematic failure modes that are obscured by prior benchmarks, yielding a more reliable and diagnostic assessment of spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。