构建开放空域三维视觉问答基准,评估大模型空间推理能力
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
- 基于真实与仿真航拍场景生成7.3万组多模态问答对
- 发现模型对相对位置判断优于绝对距离,仿真微调可提升实境表现
- 支持点云与图像模态,适合研究三维视觉语言模型的团队使用
空间推理是多模态大模型的核心能力,但其在开放空域环境中的表现仍不明确。本文提出Open3D-VQA,一个针对航拍视角下复杂空间关系推理的新基准,包含7.3万个涵盖7类任务的问答对,覆盖多项选择、是非判断和简答形式,并支持视觉与点云双模态输入。问题通过从真实与仿真航拍场景中提取的空间关系自动生成。对13个主流多模态大模型的评估显示:1)模型在相对空间关系判断上优于绝对距离理解;2)3D大模型并未显著超越2D模型;3)仅在仿真数据上微调即可显著提升模型在真实场景中的表现。我们开源了基准数据、生成流程与评估工具包:https://github.com/EmbodiedCity/Open3D-VQA.code。
原文摘要 · Abstract (English)
Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs spanning 7 general spatial reasoning tasks, including multiple-choice, true/false, and short-answer formats, and supports both visual and point cloud modalities. The questions are automatically generated from spatial relations extracted from both real-world and simulated aerial scenes. Evaluation on 13 popular MLLMs reveals that: 1) Models are generally better at answering questions about relative spatial relations than absolute distances, 2) 3D LLMs fail to demonstrate significant advantages over 2D LLMs, and 3) Fine-tuning solely on the simulated dataset can significantly improve the model's spatial reasoning performance in real-world scenarios. We release our benchmark, data generation pipeline, and evaluation toolkit to support further research: https://github.com/EmbodiedCity/Open3D-VQA.code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。