评测大模型对时空信息的精准理解能力,发现现有模型表现仍不理想。
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- 构建STI-Bench基准,涵盖机器人与车辆在多场景下的时空任务
- 顶尖大模型在距离估计和运动分析上表现不佳,精度不足
- 适合关注具身智能与自动驾驶中多模态模型评估的研究者
将多模态大语言模型(MLLMs)作为具身智能与自动驾驶的端到端解决方案已成为主流趋势。尽管MLLMs在视觉语义理解任务中已得到广泛研究,但其在真实应用中进行精确、量化的时空理解能力仍缺乏系统评估,前景不明。为此,我们提出STI-Bench,一个用于评估MLLMs时空智能的基准,涵盖桌面、室内、室外多种场景下机器人与车辆操作的挑战性任务,如物体外观、姿态、位移与运动的估计与预测。大量实验表明,当前最先进MLLMs在真实世界的时空理解任务中仍面临困难,尤其在需要精确距离估计和运动分析的任务上表现欠佳。
原文摘要 · Abstract (English)
The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their ability to perform precise and quantitative spatial-temporal understanding in real-world applications remains largely unexamined, leading to uncertain prospects. To evaluate models' Spatial-Temporal Intelligence, we introduce STI-Bench, a benchmark designed to evaluate MLLMs' spatial-temporal understanding through challenging tasks such as estimating and predicting the appearance, pose, displacement, and motion of objects. Our benchmark encompasses a wide range of robot and vehicle operations across desktop, indoor, and outdoor scenarios. The extensive experiments reveals that the state-of-the-art MLLMs still struggle in real-world spatial-temporal understanding, especially in tasks requiring precise distance estimation and motion analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。