Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
作者: Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson
分类: cs.AI
发布日期: 2026-08-21
备注: 28 pages, 2 figures
💡 一句话要点
提出DGEval基准以评估大型语言模型在IMDG合规性中的表现
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 国际海运危险品代码 合规性评估 安全关键领域 DGEval基准 决策支持工具 模型评估 运输安全
📋 核心要点
- 现有方法缺乏对大型语言模型在IMDG合规性方面的系统评估,尤其是在安全关键领域的可靠性不足。
- 论文提出DGEval基准,通过专家编写的问题和危险品清单的结构化查找,评估LLM对IMDG要求的理解能力。
- 实验结果显示,最佳模型在多项选择题上表现优于人类,但在存放、隔离和法规记忆等关键领域表现较弱,需加强人类监督。
📝 摘要(中文)
海上危险品运输是一项高风险活动,受国际海运危险品代码(IMDG Code)监管。该代码复杂,错误的分类、包装或存放可能导致严重后果。尽管大型语言模型(LLMs)被越来越多地用作决策支持工具,但尚无系统评估其在安全关键领域的可靠性。本文提出DGEval,作为评估LLM对IMDG第42-24修订版知识的基准,包含1678个问题,涵盖多种任务类型。研究表明,尽管最佳模型在多项选择题上超越人类基线,但在安全关键领域表现不佳,强调了人类监督的重要性。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在国际海运危险品代码(IMDG Code)合规性理解中的可靠性问题。现有方法未能系统评估LLM在安全关键领域的表现,导致潜在的安全隐患。
核心思路:论文提出DGEval基准,旨在通过专家编写的问题和结构化查找,系统评估LLM对IMDG第42-24修订版的知识,特别是在安全关键的操作领域。
技术框架:DGEval基准包含1678个问题,涵盖多项选择、开放式、危险品清单查找和法规识别任务。研究评估了来自六个提供商的13个模型,测试了不同思维配置的效果。
关键创新:DGEval是首个针对IMDG合规性的LLM评估基准,填补了现有研究的空白,提供了一个系统化的评估工具。与现有方法相比,DGEval专注于安全关键领域的表现,强调了人类监督的重要性。
关键设计:在实验中,模型的评估包括多种任务类型,特别关注存放、隔离和法规记忆等领域的表现。研究还测试了网络搜索对模型表现的影响,发现结构化查找与网络搜索结合能提高合规性任务的支持能力。
🖼️ 关键图片
📊 实验亮点
实验结果显示,最佳模型在多项选择题上超越人类从业者基线,但在存放、隔离和法规记忆等安全关键领域表现最弱。这表明,尽管LLMs在合规性任务中具有潜力,仍需人类监督和权威来源验证,以确保在安全关键环境中的可靠性。
🎯 应用场景
该研究的潜在应用领域包括海运危险品的合规性检查、培训和决策支持。通过提供一个系统化的评估工具,DGEval可以帮助行业从业者更好地理解和应用IMDG Code,降低运输过程中的安全风险,提升整体合规性。未来,随着模型的不断演进,DGEval可以作为持续的安全保障工具,确保LLM在安全关键领域的可靠性。
📄 摘要(原文)
The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on the NCB Hazcheck e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.