From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

📄 arXiv: 2608.09925v1 📥 PDF

作者: Laurens Samson, Iva Gornishka, Gossa Lô, Yuki M. Asano, Sennay Ghebreab

分类: cs.CL, cs.AI

发布日期: 2026-08-10

备注: Accepted at AIES 2026


💡 一句话要点

提出'Grip on LLMs'框架以评估荷兰政府用大型语言模型

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 政府应用 评估框架 荷兰 公共管理 模型选择 多语言模型

📋 核心要点

  1. 现有评估框架未能同时考虑公共管理的价值观和非英语环境的语言需求,导致政府用大型语言模型的评估不足。
  2. 提出了'Grip on LLMs'框架,通过用户研究和专家咨询,识别并操作化六个评估维度,形成系统的评估套件。
  3. 研究发现没有单一模型在所有评估维度上表现优异,且事实性与诚实性之间存在显著差异,影响模型选择的决策过程。

📝 摘要(中文)

大型语言模型在政府环境中的应用日益增加,但现有评估框架未能同时反映公共管理的价值观和非英语环境的语言需求。本文提出了'Grip on LLMs'框架,这是一个为荷兰政府使用而开发的系统评估套件,涵盖六个评估维度(事实性、诚实性、社会偏见、能耗、成本和训练数据透明度),并对30多个多语言和荷兰特定模型进行了基准测试。研究结果显示,没有单一模型在所有维度上表现优异,且高质量模型通常伴随更大的环境影响和财务成本。我们还发现,事实性和诚实性由不同属性决定,高事实性并不意味着高诚实性。为使这些发现对非技术受众可行,我们发布了一个用户友好的模型概述,供政府各方利益相关者使用。

🔬 方法详解

问题定义:本文旨在解决现有评估框架无法有效评估大型语言模型在荷兰政府环境中的应用问题,尤其是在公共管理价值观和语言需求方面的不足。

核心思路:通过与荷兰主要市政组织的领域专家合作,提出了一个系统的评估框架,涵盖多个维度,以便更全面地评估模型的适用性和性能。

技术框架:框架包括六个评估维度:事实性、诚实性、社会偏见、能耗、成本和训练数据透明度。通过用户研究、专家咨询和用户调查,形成了一个涵盖30多个模型的基准测试套件。

关键创新:最重要的创新在于识别并操作化了六个评估维度,并发现了事实性与诚实性之间的独立性,这为模型选择提供了新的视角。

关键设计:在评估过程中,采用了定量和定性的方法,结合用户反馈和专家意见,确保评估结果的全面性和实用性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,没有单一模型在所有评估维度上表现优异,且高质量模型通常伴随更大的环境影响和财务成本。具体而言,研究发现高事实性并不意味着高诚实性,这一发现为模型选择提供了新的考量标准。

🎯 应用场景

该研究的潜在应用领域包括政府部门在选择和部署大型语言模型时的决策支持,尤其是在确保模型符合公共管理价值观和语言需求方面。未来,该框架可扩展至其他非英语国家或地区的政府应用,提升公共服务的智能化水平。

📄 摘要(原文)

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.