arXiv:2608.21036cs.AI2026-08

首个评估大模型对海运危险品法规理解能力的基准,发现其在关键操作环节仍不可靠。

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

  • 构建包含1678题的DGEval基准,覆盖多类法规理解任务。
  • 最佳模型在选择题上超人类水平,但在配载、隔离等环节表现差。
  • 适合安全关键领域使用者参考,警示需人工审核与权威源验证。

海上危险品运输是高风险活动,受《国际海运危险货物规则》(IMDG)约束,错误分类、包装、配载或隔离可能导致火灾、爆炸、有毒物质泄漏或船毁人亡。正确合规需准确解读数百页相互关联的条款,且每两年更新一次。从业者越来越多地使用大语言模型(LLMs)作为决策支持工具,但尚无系统性评估证明其能否在安全关键场景中可靠应用。本文提出DGEval,首个针对IMDG第42-24号修正案的评估基准,基于专家编写的问题及危险品清单(DGL)结构化查询,涵盖1,678道多选题、开放题、DGL查找和法规识别任务。我们评估了六家厂商的13个模型,包括一个海事领域微调模型,并测试网络搜索的影响。尽管最优模型在多选题上超过人类从业者基线,但所有模型在配载、隔离和法规记忆等安全关键环节表现最弱。结果表明,LLMs可能辅助合规任务,尤其结合网络搜索进行结构化查表时,但在操作层面和法规文本召回上的不可靠性意味着部署前必须保留人工监督与权威源核验。DGEval设计为持续使用的安全保障工具,随模型演进而迭代,而非对当前能力的静态定性。

原文摘要 · Abstract (English)

The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on the NCB Hazcheck e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.

大模型评估安全合规海运规则LLM可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。