Benchmarking large language model agent societies against human behavioural distributions

📄 arXiv: 2608.28182v1 📥 PDF

作者: Raad Bin Tareaf

分类: physics.soc-ph, cs.CL

发布日期: 2026-08-28


💡 一句话要点

提出SILICA工具以评估语言模型代理社会行为

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 语言模型 社会行为 实验工具 人类行为模拟 合作动态 模型评估 人工智能伦理

📋 核心要点

  1. 现有的语言模型代理在模拟人类行为时存在不确定性,难以验证其行为是否真实反映人类社会动态。
  2. 本文提出SILICA工具,通过多种环境测试代理行为,旨在评估其与人类行为的相似性及其在不同条件下的稳定性。
  3. 实验结果显示,尽管部分模型在初始贡献上与人类相符,但在最终合作水平上存在显著差距,且模型行为受操作顺序影响明显。

📝 摘要(中文)

随着大型语言模型代理被越来越多地用作实验社会,本文提出了SILICA这一开放工具,以检验代理行为是否与人类相符、结果是否在不同实验条件下保持一致,以及社会动态是否真实存在。研究通过五个环境进行测试,结果显示,尽管部分模型在初始阶段与人类数据相符,但在最终合作状态上并未匹配。研究还发现,模型的行为受操作顺序影响显著,且大多数模型未能在协商中形成有效的共识。整体而言,当前的硅基社会仅支持探索性声明,缺乏更深层次的行为理解。

🔬 方法详解

问题定义:本文旨在解决大型语言模型代理在模拟人类行为时的真实性和一致性问题。现有方法缺乏有效的评估工具,导致结果的可靠性受到质疑。

核心思路:论文提出SILICA工具,通过设置多个实验环境并引入扰动,来检验模型的行为是否与人类行为相符,以及在不同条件下的表现是否一致。

技术框架:SILICA工具包含五个实验环境,每个环境都有与人类行为相关的基准,并通过扰动和变体来测试模型的适应性和一致性。模型在单个消费级显卡上运行,确保实验的可重复性。

关键创新:最重要的创新在于引入了扰动和变体的设计,使得模型不仅仅依赖于记忆结果,而是能够在新环境中进行适应性行为的测试。这与传统方法的单一测试环境形成鲜明对比。

关键设计:在实验中,模型的行为受到操作顺序的显著影响,且在协商过程中,只有一个模型能够正确设置接受阈值。其他模型在行为上表现出不同程度的偏差,显示出设计上的重要性。实验还表明,模型的共识形成主要依赖于共享的先验知识,而非有效的协商。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,11个模型中有8个在初始公共产品贡献上与人类数据相符,但在最终合作水平上无一匹配。仅有一个模型能够在固定的报价中正确设置接受阈值,而其他模型则表现出显著的偏差,显示出操作顺序对合作行为的影响。

🎯 应用场景

该研究的潜在应用领域包括社会科学、经济学以及人工智能伦理等。通过更准确地模拟人类行为,SILICA工具可以帮助研究人员理解社会动态,并在政策制定、市场分析等方面提供有价值的见解。未来,随着模型的不断发展,该工具的应用范围可能会进一步扩展。

📄 摘要(原文)

Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.