首个面向三维空间数值推理的多维智能评测基准,填补了3D场景下精准计算能力评估的空白。
NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities
- 构建自动化标注流程,实现多尺度精细标注与问答对生成
- 现有多模态大模型在距离、体积等精确计算上表现较差
- 适合研究三维感知、多模态推理及数值理解的学者使用
近期2D多模态大语言模型在视觉-语言任务中性能显著提升,但将其能力拓展至3D环境仍面临空间推理复杂性的挑战。现有3D基准常缺乏细粒度数值推理标注,限制了多模态大语言模型在精确空间测量和复杂数值推理方面的能力。为弥补这一差距,我们提出NUMINA——首个面向多维智能与数值推理能力的自然理解基准,旨在提升多模态室内感知理解能力。NUMINA采用多尺度标注和多样问答对,通过NUMINA-Flow自动化标注流程生成,该流程融合大语言模型重写与基于规则的自验证机制。我们基于Chat-Scene框架评估多种先进大语言模型在NUMINA上的表现,结果表明当前模型在多模态数值推理方面存在明显不足,尤其在距离与体积估算等精确计算任务上表现不佳,凸显3D模型进一步发展的必要性。数据集与源代码可从https://github.com/fengshun124/NUMINA获取。
原文摘要 · Abstract (English)
Recent advancements in 2D multimodal large language models (MLLMs) have significantly improved performance in vision-language tasks. However, extending these capabilities to 3D environments remains a distinct challenge due to the complexity of spatial reasoning. Nevertheless, existing 3D benchmarks often lack fine-grained numerical reasoning task annotations, limiting MLLMs' ability to perform precise spatial measurements and complex numerical reasoning. To address this gap, we introduce NUMINA, the first Natural Understanding benchmark for Multi-dimensional Intelligence and Numerical reasoning Abilities to enhance multimodal indoor perceptual understanding. NUMINA features multi-scale annotations and various question-answer pairs, generated using NUMINA-Flow, an automated annotation pipeline that integrates LLM rewriting and rule-based self-verification. We evaluate the performance of various state-of-the-art LLMs on NUMINA following the Chat-Scene framework, demonstrating that current LLMs struggle with multimodal numerical reasoning, particularly in performing precise computations such as distance and volume estimation, highlighting the need for further advancements in 3D models. The dataset and source codes can be obtained from https://github.com/fengshun124/NUMINA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。