Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

📄 arXiv: 2608.12278v1 📥 PDF

作者: Avijit Roy, Proma Roy

分类: cs.CL, cs.AI, cs.CY

发布日期: 2026-08-12

备注: An associated poster version of this work was presented at the 69th Annual Conference of the International Linguistic Association (ILA 2026), New York, NY, April 30-May 2, 2026


💡 一句话要点

探讨AI基础设施对弱势语言使用者的影响与解决方案

🎯 匹配领域: 支柱四:生成式动作 (Generative Motion)

关键词: 弱势语言支持 AI基础设施 教育技术 数据集稀缺 离线优先设计 社会公平 语言学研究

📋 核心要点

  1. 现有AI工具在教育和语言支持方面未能有效服务于弱势语言使用者,尤其是在基础设施设计上存在系统性缺陷。
  2. 论文提出将数据集稀缺视为结构性障碍,并倡导离线优先设计作为解决方案,以提升弱势语言的AI支持。
  3. 通过分析孟加拉语的具体案例,识别出四个主要的结构性失败,强调了对弱势语言的关注不足。

📝 摘要(中文)

人工智能教育和语言支持工具被视为解决资源匮乏社区的可扩展响应。然而,这些工具背后的基础设施,包括训练语料、分词方案、评估基准和部署架构,可能在模型训练前就系统性地使弱势语言使用者处于不利地位。本文以孟加拉语为例,分析了四个相互关联的失败,包括网络内容缺口、训练标记不足、分词惩罚和连接排斥,反映了长期以来的资源分配决策和设计缺陷。我们认为数据集稀缺应被视为结构性障碍,而离线优先设计应作为一种公平导向的基础设施策略。最后,我们提出了减少这些结构性不平等的语言学和AI研究方向。

🔬 方法详解

问题定义:本文旨在解决AI基础设施对弱势语言使用者的系统性不利影响,现有方法未能考虑这些语言的特殊需求,导致教育和语言支持工具的有效性受限。

核心思路:论文提出将数据集稀缺视为结构性障碍,强调在AI开发中需要关注弱势语言的资源配置,并倡导离线优先的设计理念,以确保这些语言的使用者能够获得平等的技术支持。

技术框架:整体架构包括对现有数据集的分析、对比不同语言的训练标记数量、评估分词效果以及研究网络连接的可及性,主要模块包括数据收集、分析和设计优化。

关键创新:最重要的创新点在于将数据集稀缺视为结构性问题,而非单纯的技术限制,强调了在AI工具设计中需要考虑弱势语言的特殊需求。

关键设计:在设计过程中,关注了训练标记的数量、分词算法的适应性以及网络连接的可及性,确保在低连接环境中也能有效支持弱势语言的学习和使用。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,孟加拉语在主要多语言语料中的训练标记数量与英语相比存在67:1的差距,且网络内容占比不足0.5%。这些数据突显了弱势语言在AI支持中的严重不足,呼吁对其进行更深入的关注和资源投入。

🎯 应用场景

该研究的潜在应用领域包括教育技术、语言学习应用和社会公平技术。通过改善对弱势语言的支持,可以促进这些语言的传播和使用,提升其在全球化背景下的可见性和价值,进而推动社会的公平与包容。

📄 摘要(原文)

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.