首个专用于海图理解的多模态大模型评测基准,检验AI能否读懂专业航海图。
ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
- 构建三层次海图理解任务:符号识别、空间推理与航行决策
- 10个顶尖模型在2万+专家标注样本上平均仅47.88%准确率
- 揭示模型在符号对齐、空间计算和多约束推理上的系统性缺陷
电子海图(ENCs)是现代航海安全的核心,但多模态大语言模型(MLLMs)是否能可靠解读仍不明确。与自然图像或常规地图不同,ENCs通过标准化矢量符号、比例依赖渲染和精确几何结构编码法规、水深及航线限制,需专门航海知识才能理解。本文提出ENC-Bench,首个专注于专业海图理解的基准。该数据集包含来自840份美国国家海洋和大气管理局(NOAA)ENC的20,490个专家验证样本,按三层次组织:感知(符号与特征识别)、空间推理(坐标定位、方位角、距离)和航海决策(航线合法性、安全评估、多重约束下的应急规划)。所有样本均通过校准的矢量转图像管道从原始S-57数据生成,并经自动化一致性检查与专家审核。我们在统一零样本协议下评估了10个前沿MLLMs,包括GPT-4o、Gemini 2.5、Qwen3-VL、InternVL-3和GLM-4.5V。最佳模型仅达47.88%准确率,暴露出符号对齐、空间计算、多约束推理及光照/尺度变化鲁棒性等方面的系统性挑战。本研究建立了首个严谨的海图理解基准,开辟了专用符号推理与安全关键AI交叉的新方向,为提升MLLM在专业航海应用中的能力提供了关键基础设施。
原文摘要 · Abstract (English)
Electronic Navigational Charts (ENCs) are the safety-critical backbone of modern maritime navigation, yet it remains unclear whether multimodal large language models (MLLMs) can reliably interpret them. Unlike natural images or conventional charts, ENCs encode regulations, bathymetry, and route constraints via standardized vector symbols, scale-dependent rendering, and precise geometric structure -- requiring specialized maritime expertise for interpretation. We introduce ENC-Bench, the first benchmark dedicated to professional ENC understanding. ENC-Bench contains 20,490 expert-validated samples from 840 authentic National Oceanic and Atmospheric Administration (NOAA) ENCs, organized into a three-level hierarchy: Perception (symbol and feature recognition), Spatial Reasoning (coordinate localization, bearing, distance), and Maritime Decision-Making (route legality, safety assessment, emergency planning under multiple constraints). All samples are generated from raw S-57 data through a calibrated vector-to-image pipeline with automated consistency checks and expert review. We evaluate 10 state-of-the-art MLLMs such as GPT-4o, Gemini 2.5, Qwen3-VL, InternVL-3, and GLM-4.5V, under a unified zero-shot protocol. The best model achieves only 47.88% accuracy, with systematic challenges in symbolic grounding, spatial computation, multi-constraint reasoning, and robustness to lighting and scale variations. By establishing the first rigorous ENC benchmark, we open a new research frontier at the intersection of specialized symbolic reasoning and safety-critical AI, providing essential infrastructure for advancing MLLMs toward professional maritime applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。