MapTab用地图+表格测试模型多条件路线规划能力
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs
- 用地图图像与表格数据结合,评估模型多条件路线推理
- 涵盖328张图、近20万条查询,含时间、价格等四维度指标
- 适合研究多模态推理、智能导航的学者与开发者
系统评估多模态大语言模型(MLLMs)对推动通用人工智能(AGI)至关重要。然而现有基准难以严格衡量其在多条件约束下的推理能力。为此,我们提出MapTab,一个通过路线规划任务评估MLLMs综合多条件推理能力的多模态基准。MapTab要求模型从地图图像中感知并定位视觉信息,同时整合结构化表格中的路线属性,如时间与价格。包含两种场景:Metromap覆盖全球52个国家160座城市的地铁网络,Travelmap包含19个国家168个代表性景点。整体包含328张图像、196,800条路线规划查询和3,936个问答查询,涵盖时间、价格、舒适度和可靠性四个关键标准。对21个代表性MLLMs的广泛评估显示,当前模型在多条件多模态推理方面仍存在明显不足。尤其当视觉感知不可靠时,多模态推理甚至可能劣于单模态方法。MapTab因此提供了一个具有挑战性且贴近现实的测试平台,用于系统评估与推进MLLMs在核心感知、信息融合、数值比较及路线规划等方面的能力。
原文摘要 · Abstract (English)
Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for rigorously measuring their reasoning capabilities under multi-criteria constraints. To address this gap, we introduce MapTab, a multimodal benchmark designed to assess holistic multi-criteria reasoning in MLLMs through route-planning tasks. MapTab requires models to perceive and ground visual information from map images while integrating route attributes, such as Time and Price, from structured tables. It covers two scenarios: Metromap, spanning metro networks in 160 cities across 52 countries, and Travelmap, featuring 168 representative tourist attractions from 19 countries. Overall, MapTab includes 328 images, 196,800 route-planning queries, and 3,936 QA queries, incorporating four key criteria: Time, Price, Comfort, and Reliability. Extensive evaluations of 21 representative MLLMs show that current models still struggle with multicriteria multimodal reasoning. Notably, when visual perception is unreliable, multimodal reasoning can even underperform unimodal approaches. MapTab therefore offers a challenging and realistic testbed for systematically evaluating and advancing MLLMs across core perception, integration, numerical comparison, and route planning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。