评测大模型在室内视频中精确定位小物体的能力,填补了空间理解评估的空白。
PinpointQA: A Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos

- 基于ScanNet++和ScanNet200构建,包含1024个场景与10094组问答对。
- 四项渐进任务中结构化空间预测准确率最低,表明模型定位能力仍有不足。
- 专为诊断和训练多模态模型的空间定位能力设计,适合研究具身智能者使用。
在室内环境中实现可靠的具身交互需要智能体从视觉观测中精确定位小型日常物品。然而,这一基础能力对多模态大语言模型(MLLMs)仍具挑战性,尤其是在从室内视频中进行空间理解时。现有基准侧重于视频空间智能与具身推理,但未直接评估模型对小目标物体的精确定位及位置表达能力。我们提出PinpointQA,基于ScanNet++和ScanNet200构建,包含1,024个场景和10,094组问答对,涵盖四项逐步递增难度的任务:目标存在性验证、最近参考物识别、细粒度空间描述与结构化空间预测。真值标注源自对齐的3D几何与实例级标注的中间空间表示,而评估模型仅接收采样的RGB视频帧。对代表性MLLMs的评估显示,性能随任务难度递增持续下降,其中结构化空间预测尤为困难;监督微调显著提升表现。通过分离下游具身交互前的空间定位能力,PinpointQA兼具诊断与训练价值。
原文摘要 · Abstract (English)
Reliable embodied interaction in indoor environments requires agents to precisely localize small everyday objects from visual observations. Yet this fundamental capability remains challenging for multimodal large language models (MLLMs), particularly when spatial understanding must be performed from indoor videos. Existing benchmarks study video spatial intelligence and embodied reasoning, but do not directly evaluate whether a model can localize a small target object and express its position with sufficient precision. We introduce PinpointQA, a benchmark built from ScanNet++ and ScanNet200. It contains 1,024 scenes and 10,094 QA pairs across four progressively challenging tasks: Target Presence Verification, Nearest Reference Identification, Fine-Grained Spatial Description, and Structured Spatial Prediction. Ground-truth annotations are constructed from intermediate spatial representations derived from aligned 3D geometry and instance-level annotations, while evaluated models receive only sampled RGB video frames. Evaluations of representative MLLMs reveal a consistent performance decline across the task progression, with structured spatial prediction remaining particularly challenging, while supervised fine-tuning substantially improves performance. By isolating the spatial grounding capabilities that precede downstream embodied interaction, PinpointQA serves as both a diagnostic benchmark and an effective training resource. The project page is available at https://rainchowz.github.io/PinpointQA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。