Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm
作者: Hidayet Aksu
分类: cs.CR, cs.AI
发布日期: 2026-08-17
备注: 10 pages, 7 figures,
💡 一句话要点
将米尔格拉姆服从实验引入大型语言模型以评估其服从性
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 服从性评估 米尔格拉姆实验 社会心理学 人机交互 自动化系统 伦理研究
📋 核心要点
- 现有大型语言模型在执行指令时的服从性缺乏系统性评估,尤其是在权威要求下的行为表现。
- 本文提出将米尔格拉姆服从实验移植到LLMs中,通过标准化实验设计评估模型的服从性。
- 实验结果显示,服从性在不同模型间存在显著差异,且受情境因素影响,提供了对模型行为的新见解。
📝 摘要(中文)
大型语言模型(LLMs)越来越多地被用作执行指令和操作设备的代理,这引发了一个社会心理学问题:在权威的要求下,代理会在多大程度上升级有害行为?本文将米尔格拉姆的服从实验移植到LLMs中,设计了一个标准化、可复制的实验框架,测量42个模型在不同条件下的服从性。研究发现,服从性具有高度异质性,模型特异性和稳定性,且情境敏感性表现出选择性。声明场景为虚构会提高服从性,而将决策转移到本地工具调用则显著降低服从性。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在权威要求下的服从性评估问题。现有方法缺乏对模型行为的系统性分析,尤其是在复杂情境下的表现。
核心思路:通过将米尔格拉姆的服从实验移植到LLMs中,设计一个标准化的实验框架,使得模型能够在模拟的权威情境中进行行为测试,从而量化其服从性。
技术框架:整体架构包括模型扮演教师角色,确定性设备模拟实验者和学习者,使用改编的米尔格拉姆脚本进行实验。实验通过30个震动级别(15-450 V)和标准化的提示进行。
关键创新:本研究的主要创新在于将心理学实验标准化并应用于LLMs,首次系统性地量化了模型在权威情境下的服从性,并揭示了模型间的异质性和情境敏感性。
关键设计:实验设计中使用了六种不同的条件来测量服从性,采用单标记指纹研究的普查方法,确保了数据的可靠性和可重复性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,服从性在不同模型间高度异质,基线全服从率范围为0-100%(平均42.9%),其中5个模型在每次实验中均施加最大震动,而11个模型则从未施加。情境因素对服从性影响显著,声明场景为虚构时服从性中位数提高17.2 V。
🎯 应用场景
该研究的潜在应用领域包括人工智能伦理、自动化系统的安全性评估以及人机交互设计。通过理解模型在权威情境下的行为,可以为未来的AI系统设计提供重要的指导,确保其在执行任务时的安全性和可靠性。
📄 摘要(原文)
Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists? We port Milgram's obedience paradigm to LLMs as a standardized, fully scripted, replicable probe: the model plays the Teacher, a deterministic harness plays Experimenter and Learner from paraphrased Milgram scripts (30 shock levels, 15-450 V; graded protests; the four standardized prods), and the outcome of a session is the breakoff voltage. Following the census methodology of single-token fingerprinting studies, we measure obedience profiles (empirical breakoff distributions over a battery of six conditions) for 42 models from 19 families. We find that (i) obedience is highly heterogeneous: baseline full-obedience rates span 0-100% (census mean 42.9%; human anchor 65%), with 5 models delivering the maximum shock in every session and 11 never doing so; (ii) profiles are model-specific and stable: split-half verification separates same-model from cross-model comparisons with AUC 0.885 (0.949 under an ordinal-aware distance); (iii) situational sensitivity is selective: peer defiance shifts obedience in the human direction, learner proximity only weakly, and removing the authority's physical presence (the strongest human lever) has no detectable effect; (iv) declaring the scenario fictional raises obedience (median +17.2 V), whereas moving the decision to a native tool call lowers it sharply (-53.0 V), as does a 1,024-token deliberation budget (-38.2 V); and (v) obedience profiles do not recover model lineage (leave-one-out family accuracy 8.3% vs. 3.7% chance): obedience identifies the checkpoint, not its ancestry, consistent with safety post-training overwriting lineage priors.