BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian
作者: Jophin John, Michael Hoffmann, Jan Fillies, Michael A. Hedderich, Barbara Plank
分类: cs.CL
发布日期: 2026-08-13
💡 一句话要点
提出BavGround基准以评估巴伐利亚地区文化与方言能力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 文化评估 方言能力 多语种评估 巴伐利亚文化 基准测试 自然语言处理
📋 核心要点
- 现有的大型语言模型评估往往忽视地区文化和方言,导致这些领域的知识缺乏代表性。
- 论文提出BavGround基准,通过多项选择题评估巴伐利亚地区文化和方言能力,涵盖多种语言。
- 实验结果表明,多语种模型在整体表现上优于其他模型,但在方言和地方文化知识上仍存在显著不足。
📝 摘要(中文)
本研究聚焦于大型语言模型(LLMs)的文化评估,指出现有方法多集中于高资源标准语言,导致地区文化和方言社区的代表性不足。我们引入BavGround,一个用于评估巴伐利亚地区文化基础和方言能力的基准,涵盖英语、德语和巴伐利亚语。BavGround包含206个多项选择题,跨越八个文化领域,提供618个多语种实例,涵盖广泛的文化知识和源自新闻、历史文献及专业文献的地区知识。对15个7B-10B的开放权重指令调优模型和一个闭源模型的评估显示,尽管强大的多语种模型表现最佳,但在巴伐利亚项目和源基础问题上的表现下降,表明方言和地方文化知识的掌握仍然存在困难。
🔬 方法详解
问题定义:本研究旨在解决大型语言模型在地区文化和方言知识评估中的不足,现有方法多集中于高资源语言,忽视了地方文化的多样性和复杂性。
核心思路:BavGround基准通过设计多项选择题,涵盖广泛的文化领域和地方知识,提供了一种系统的评估方式,以检测模型在地区文化和方言能力上的表现。
技术框架:BavGround的整体架构包括题库的构建、模型评估和结果分析三个主要模块。题库包含206个问题,跨越八个文化领域,模型评估则使用多种评分协议进行。
关键创新:BavGround的主要创新在于其多语种和多文化的评估框架,特别是对方言和地区文化知识的关注,这与传统的评估方法形成了鲜明对比。
关键设计:在设计中,采用了多种评分协议,包括原始答案评分、选项文本可能性和语义匹配等,以确保评估的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,强大的多语种模型在整体评估中表现最佳,但在巴伐利亚项目和源基础问题上的表现显著下降,表明方言能力仍然较弱。不同的评估协议导致的评分差异,强调了评估方法选择的重要性。
🎯 应用场景
BavGround基准的提出为大型语言模型在地区文化和方言能力的评估提供了新的工具,具有广泛的应用潜力。它可以用于改进语言模型的训练,使其更好地理解和生成地方文化相关内容,提升模型在多样化语言环境中的表现,具有重要的社会和文化价值。
📄 摘要(原文)
Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian regional cultural grounding and dialect competence across English, German and Bavarian. BavGround contains 206 multiple-choice source questions across eight cultural domains per language, yielding 618 multi-parallel instances, with items covering both broadly accessible cultural knowledge and source-grounded regional knowledge from journalism, historical sources, and specialist literature. We evaluate fifteen 7B-10B open-weight instruction-tuned models and one closed-model reference. Strong multilingual models perform best overall, but performance drops on Bavarian items and source-grounded questions, indicating persistent difficulty with dialectal and localized cultural knowledge. We further show that conclusions depend strongly on evaluation protocol: raw answer-letter scoring, shuffled-letter scoring, option-text likelihood, generated-answer parsing, and semantic matching can produce different absolute scores and rankings, especially for regionally adapted models. Finally, an exploratory analysis of GENBA-10B checkpoints suggests that continued pretraining improves answer-content likelihood unevenly across domains, while dialect competence remains comparatively weak. BavGround supports localized, protocol-aware evaluation of cultural representation in LLMs.