Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making
作者: Jiayuan Di, Haoyi Yang, Yufei Luo, Jiahui Qu, Yiming Wang
分类: cs.CL
发布日期: 2026-07-29
💡 一句话要点
提出S2D框架以系统评估大型语言模型中的区域偏见
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 区域偏见 大型语言模型 刻板印象 社会决策 评估框架 人工智能伦理 中国研究
📋 核心要点
- 现有研究往往将区域偏见的表现分开考察,导致对其结构和后果的理解不够全面。
- 本文提出S2D框架,系统评估从抽象刻板印象到具体社会决策的区域偏见,涵盖多个维度。
- 实验结果显示,区域评分存在显著差异,尤其在能力和职业决策上,各模型间高度一致,且与区域发展指标相关联。
📝 摘要(中文)
大型语言模型(LLMs)中的区域偏见可能影响对不同地区群体的看法以及对个体的决策。然而,现有研究往往将这些表现分开考察,导致其结构和后果不明确。本文提出了Stereotypes-to-Decisions(S2D)框架,系统评估从抽象刻板印象到具体社会决策的区域偏见。覆盖中国34个省级行政区域,S2D通过对温暖度(友好和可信度)和能力(能力和智力)的刻板印象评分,以及教育、职业和社会互动的配对选择任务,评估六个LLMs。结果显示,区域评分存在显著差异,各模型间在能力和职业决策上高度一致。此外,这些模式与区域经济和数字发展指标相关联,显示出混合的人类刻板印象。总体而言,研究表明LLMs中的区域偏见普遍、系统且具有后果,促使更具区域意识的评估与缓解。
🔬 方法详解
问题定义:本文旨在解决大型语言模型中区域偏见的评估问题,现有方法未能全面考察其对社会决策的影响。
核心思路:提出Stereotypes-to-Decisions(S2D)框架,通过系统化的方式将抽象刻板印象与具体决策联系起来,提供更全面的评估视角。
技术框架:S2D框架包括刻板印象评分和配对选择任务两个主要模块,前者评估温暖度和能力,后者涉及教育、职业和社会互动的决策。
关键创新:该框架的创新在于将区域偏见的评估从抽象层面提升到具体决策层面,填补了现有研究的空白。
关键设计:在实验中,采用了多种LLMs进行评估,并结合区域经济和数字发展指标,确保结果的可靠性和有效性。具体参数设置和损失函数设计未在摘要中详细说明,需参考完整论文。
🖼️ 关键图片
📊 实验亮点
实验结果显示,六个大型语言模型在能力和职业决策上表现出高度一致性,区域评分存在显著差异,尤其在温暖度和能力维度上。研究还发现这些模式与区域经济和数字发展指标相关,显示出区域偏见的系统性和后果。
🎯 应用场景
该研究的潜在应用领域包括社会科学研究、政策制定和人工智能伦理。通过识别和评估区域偏见,能够帮助改进大型语言模型的公平性和透明度,促进更具包容性的技术发展。未来,S2D框架可扩展至其他国家和文化背景,推动全球范围内的区域偏见研究。
📄 摘要(原文)
Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions. Covering all 34 provincial-level administrative regions of China, S2D evaluates six LLMs using stereotype ratings of Warmth (perceived friendliness and trustworthiness) and Competence (perceived capability and intelligence), along with paired-choice tasks across Education, Occupation, and Social Interaction. Results reveal substantial regional differences in regional scores, with considerable agreement across models, especially for Competence and Occupation decisions. Furthermore, these patterns are associated with regional economic and digital development indicators and display mixed human-like stereotypes, with some regions rated highly on one dimension but poorly on the other. They also remain largely stable across Chinese and English prompts. Overall, our findings show that regional bias in LLMs is prevalent, systematic, and consequential, motivating more regionally aware evaluation and mitigation.