The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations
作者: Sandra C. Sandoval, Navita Goyal, Rashawn Ray, Long Doan, Rachel Rudinger, Hal Daumé
分类: cs.CY, cs.AI, cs.CL
发布日期: 2026-08-05
💡 一句话要点
探讨种族与性别对警察语言使用的影响
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 警察语言使用 种族偏见 虚拟现实 因果推断 大型语言模型 社会影响 警民关系
📋 核心要点
- 现有研究未能充分揭示警察在与不同种族和性别角色互动时的语言差异,尤其是在高压情境下。
- 本研究通过虚拟现实模拟,采用因果推断方法,分析警察与黑人男性角色的交流,评估语言使用的尊重程度。
- 实验结果显示,警察对黑人男性角色的语言尊重程度显著低于其他角色,且这种差异在特定情境下更为明显。
📝 摘要(中文)
在美国警察与公众互动暴力的背景下,本研究探讨警察在虚拟现实(VR)模拟中与被描绘为黑人男性的虚拟角色交流时的语言使用差异。通过因果推断方法,我们发现大多数警察对黑人男性角色的语言表现出较低的尊重程度,尤其是在角色被视为嫌疑犯的情况下。研究还探讨了大型语言模型(LLMs)在平均处理效应(ATE)估计中的应用,提出了混合效应模型与逆倾向加权(iptw)方法的结合,展示了LLM辅助方法在此领域的潜力。该研究为理解警察语言使用的社会影响提供了重要见解,并为未来的研究指明了方向。
🔬 方法详解
问题定义:本研究旨在解决警察在与不同种族和性别角色互动时语言使用的差异问题。现有方法未能充分捕捉这种差异的社会影响,尤其是在高压情境下的交流。
核心思路:通过虚拟现实(VR)模拟,研究警察与黑人男性角色的互动,采用因果推断方法评估语言的尊重程度,揭示种族和性别对警察语言使用的影响。
技术框架:研究设计包括VR模拟环境的构建、角色分配、数据收集与分析。主要模块包括角色特征设定、警察语言分析和因果推断模型的应用。
关键创新:本研究的创新点在于结合虚拟现实技术与因果推断方法,首次系统性地分析警察在不同种族角色下的语言使用差异,提供了新的视角和实证数据。
关键设计:采用混合效应模型与逆倾向加权(iptw)方法进行ATE估计,利用大型语言模型(LLM)进行文本特征创建,确保模型的准确性与可靠性。
🖼️ 关键图片
📊 实验亮点
实验结果表明,大多数警察对黑人男性角色的语言尊重程度显著低于其他角色,尤其是在嫌疑犯情境下,尊重程度差异可达2至数个点(0-10分制)。此外,研究还展示了LLM在ATE估计中的潜力,为未来研究提供了新的方法论基础。
🎯 应用场景
该研究的结果对警务培训、政策制定和社会心理学研究具有重要的应用价值。通过理解警察在不同种族角色下的语言使用差异,可以为改善警民关系、减少误解和暴力提供理论支持。此外,研究方法的创新为其他社会科学领域的因果推断提供了新的思路。
📄 摘要(原文)
Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through a causal in- ference lens, where the assignment of the Black man character to a police officer and simulation is the treatment variable. Our (marginal) average treatment effect AT E measures the social impact of the character on the deference of officer statements with each turn of the conversation. Soberingly, we find that most officers speak less deferentially to Black man characters, except for White, biracial, and multiracial female officers, es- pecially in settings where the VR character was known to be a suspect. Across a full conversation of a typical VR scene, these marginal AT Es can result in notable changes in def- erence of tone (two to several points difference on a scale of 0-10), above and beyond that due to the initial effect of per- ceiving a Black male character. Even more disconcerting is that this can contribute to conversation breakdowns that po- tentially result in violence or danger to both the public and the police. We also explored the capabilities of large language models (LLMs) for ATE estimation. From our methods com- parison analysis, including model validation against synthetic data, we provide unique scientific insights on LLM-assisted methodologies for ATE estimation. As such, for ATE esti- mation with multilevel data with text, we recommend mixed effects models with the inverse propensity treatment weighted (iptw) approach, which utilized an LLM for text feature cre- ation. While we also tested LLMs for finetuning prediction models ultimately for ATE estimation, we conclude they are an area for further development and refinement.