构建多模态空间认知基准,评估模型对物体属性与空间关系的细粒度理解能力。
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
- 设计三类任务:计数、空间关系推理、带关系的计数
- 现有模型在细粒度识别和动态推理上表现不足,尤其在复杂场景中
- 适合研究空间感知、视觉推理与多模态智能的学者使用
空间感知与推理是人类认知的核心,涵盖物体识别、空间关系理解及动态推理。尽管计算机视觉取得进展,现有基准仍暴露出模型在准确识别物体属性和推理空间关系方面的显著缺陷,而这正是动态推理的基础。为此,我们提出 MIRAGE,一个面向多模态空间感知、推理与智能的基准,包含计数(物体属性识别)、关系(空间关系推理)以及带关系的计数三类任务。通过多样化且复杂的场景,MIRAGE揭示了当前最先进模型在细粒度识别与推理上的关键局限,凸显改进表示学习与推理框架的必要性。该基准为未来时空推理研究提供了重要路径。
原文摘要 · Abstract (English)
Spatial perception and reasoning are core components of human cognition, encompassing object recognition, spatial relational understanding, and dynamic reasoning. Despite progress in computer vision, existing benchmarks reveal significant gaps in models' abilities to accurately recognize object attributes and reason about spatial relationships, both essential for dynamic reasoning. To address these limitations, we propose MIRAGE, a multi-modal benchmark designed to evaluate models' capabilities in Counting (object attribute recognition), Relation (spatial relational reasoning), and Counting with Relation. Through diverse and complex scenarios requiring fine-grained recognition and reasoning, MIRAGE highlights critical limitations in state-of-the-art models, underscoring the need for improved representations and reasoning frameworks. By targeting these foundational abilities, MIRAGE provides a pathway toward spatiotemporal reasoning in future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。