AI systems and the reproduction of (standard) language ideologies in World Englishes
作者: Kingsley Ugwuanyi
分类: cs.CL
发布日期: 2026-07-30
备注: 13 pages, 0 figure
💡 一句话要点
探讨AI系统如何重现语言意识形态以解决英语标准化问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 语言意识形态 大型语言模型 社会语言学 英语标准化 多样性 AI技术 全球南方英语 算法偏见
📋 核心要点
- 核心问题:AI系统在语言模型训练中重现了主导的语言意识形态,导致非主流英语被边缘化。
- 方法要点:通过实证研究和媒体分析,探讨AI如何在不同层面上影响语言标准化和合法性。
- 实验或效果:分析表明,AI技术在促进标准化的同时,也可能通过多样化语料库来引发对英语多样性的关注。
📝 摘要(中文)
随着大型语言模型(LLMs)的快速发展,社会语言学和世界英语的传统问题再次浮现,例如谁决定什么算是合法英语,谁的英语被视为可疑等。本文考察了AI系统及其使用和相关话语如何反映、强化并偶尔挑战(标准)语言意识形态,这些意识形态优先考虑内圈规范并边缘化非主流英语。通过实证研究、媒体评论和社交媒体辩论的证据,本文展示了AI技术在不同层面上重现主导语言意识形态,包括训练数据、设计协议、评估基准、用户反馈和公众评论。分析利用了围绕AI语言的公众争议,特别是对“delve”一词的关注,说明全球北方的英语使用者如何监管全球南方英语使用者的语言规范。本文还指出了Christian Mair所称的“标准化悖论”,即AI可能通过优先考虑标准形式来同质化英语,同时通过全球南方用户的广泛语料和注释工作来多元化英语。
🔬 方法详解
问题定义:本文旨在解决AI系统如何在语言模型中重现和强化主导语言意识形态的问题。现有方法未能充分考虑非主流英语的合法性,导致其在AI生成内容中被忽视或边缘化。
核心思路:论文通过分析AI系统的设计和使用,提出需要更具包容性的设计方法,以承认英语的多样性,从而减少对非主流英语的偏见。
技术框架:整体架构包括数据收集、模型训练、用户反馈和公众评论分析等多个阶段。每个阶段都涉及对语言意识形态的反思和评估。
关键创新:最重要的技术创新在于识别和分析AI生成内容中潜在的语言意识形态,尤其是如何通过算法和训练数据影响语言标准化。与现有方法的区别在于关注了全球南方用户的贡献和视角。
关键设计:在设计过程中,考虑了多样化的语料库和注释工作,确保模型训练中包含不同英语变体的代表性。此外,评估标准也应反映多样性,而非单一的标准形式。
🖼️ 关键图片
📊 实验亮点
研究表明,AI技术在语言生成中重现了主导的语言意识形态,尤其是在对特定词汇的使用上,导致非主流英语的边缘化。通过对比分析,发现AI生成的内容在标准化和多样性之间存在复杂的张力。
🎯 应用场景
该研究的潜在应用领域包括教育、语言政策和AI系统设计。通过更好地理解AI如何影响语言意识形态,可以促进对多样性和包容性的重视,从而在全球化背景下更好地使用和理解英语。
📄 摘要(原文)
The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc. This paper examines how AI systems, their uses and discourse on them reflect, reinforce, and occasionally challenge (standard) language ideologies, which privilege Inner Circle norms and marginalize non-dominant Englishes. Drawing on evidence from empirical studies, media commentary, social media debates, and examples from AI outputs, the paper shows that AI technologies reproduce dominant language ideologies at different levels: training data, design protocols, evaluation benchmarks, user feedback and public commentary. The analysis uses the public controversy over AI-sounding language, especially the fixation on the word delve, to illustrate how speakers of English from the Global North police the English language norms of Global South English users. The paper also identifies what Christian Mair has called a "standardisation paradox": AI may homogenize English by privileging standard forms and at the same time pluralize Englishes through exposure to wide-ranging corpora and annotation work carried out by Global South users. In doing so, the paper argues that generative AI is reigniting long-standing debates in World Englishes about standardization, legitimacy, and the ownership of English, now playing out in algorithmic systems, model training, evaluation practices, and public discourse, where non-dominant Englishes are increasingly conflated with AI-generated speech. Discussing AI systems as a site where language ideologies are (re)produced, the paper argues for more inclusive design approaches that recognize the plurality of Englishes in order to address the real-world negative consequences of treating some as more legitimate than others.