Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

📄 arXiv: 2608.20438v1 📥 PDF

作者: Rana Muhammad Usman, Dominic Williamson

分类: physics.soc-ph, cs.AI, cs.MA, cs.SI

发布日期: 2026-08-20

备注: 9 pages, 3 figures. Code, frozen protocol, configurations, summary tables, and data are publicly available at https://github.com/ranausmanai/synthetic-social-networks and https://huggingface.co/datasets/ranausmans/synthetic-social-networks


💡 一句话要点

提出PV-SST以评估LLM代理的词汇收敛性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 词汇收敛 同行投票 社交平台 实验设计

📋 核心要点

  1. 现有方法无法通过单一代理基准全面表征大型语言模型代理的群体行为,导致对其性能的理解不足。
  2. 本文提出PV-SST测试平台,通过同行投票机制评估LLM代理在不同主题下的表现,旨在揭示词汇收敛现象。
  3. 实验结果表明,经过同行排名的帖子馈送显著提高了词汇相似性,但未能证明分布式来源在诚实代理立场上的优势。

📝 摘要(中文)

在大型语言模型(LLM)代理的研究中,单一代理基准无法全面表征群体行为。本文提出了PV-SST,一个基于同行投票的社交平台测试环境,并报告了一项涵盖四个主题、四个未使用种子、四个开放权重模型系列及三个预设更大变体的匹配暴露实验。实验结果显示,相较于仅基于主题的控制组,经过同行生成点赞排名的前一轮帖子馈送显著提高了最终轮的词汇相似性。研究发现,尽管在核心面板中对立生存率有所下降,但在更大变体中并未得出明确结论。总体而言,研究结果表明,经过同行排名的馈送促进了词汇收敛,而非一般的意见捕获或协调优势。

🔬 方法详解

问题定义:本文旨在解决大型语言模型(LLM)代理的群体行为无法通过单一代理基准全面表征的问题。现有方法未能有效捕捉到多代理互动中的复杂性与动态变化。

核心思路:论文提出了PV-SST,一个基于同行投票的社交平台测试环境,旨在通过多轮实验评估LLM代理在不同主题下的表现,特别关注词汇收敛现象。通过引入同行生成的点赞排名,研究探讨了信息馈送对模型表现的影响。

技术框架:整体架构包括四个主题、四个未使用的随机种子和四个开放权重模型系列,实验设计涵盖448次试验和112个完整的模型-主题-种子块。每个实验块中,模型通过接收前一轮的同行帖子进行训练和评估。

关键创新:最重要的技术创新在于引入了同行投票机制和排名系统,研究发现这种设计能够有效促进词汇收敛,而非单纯依赖于信息的多样性或分布式来源。

关键设计:实验中采用了TF-IDF余弦相似度作为评估指标,设置了多个对照组以确保结果的可靠性。模型的训练和评估过程严格按照预注册的实验设计进行,确保了实验的可重复性和透明性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,经过同行生成点赞排名的帖子馈送显著提高了最终轮的词汇相似性,核心面板的配对均值差异为+0.0082 TF-IDF余弦单位,且具有统计显著性(p=0.000105)。在更大变体中,词汇相似性提升更为明显,达到+0.0109(p=0.000001),显示出同行排名对模型表现的积极影响。

🎯 应用场景

该研究的潜在应用领域包括社交媒体内容生成、在线教育平台和人机交互系统等。通过理解LLM代理的群体行为,开发者可以优化模型的训练策略,提高其在多用户环境中的表现,进而提升用户体验和内容质量。未来,研究结果可能为LLM的协作学习和信息传播机制提供新的视角。

📄 摘要(原文)

Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.