How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
作者: Hasan Mahmud, Khawaja Abaid Ullah, Mohammad Javad Khojasteh, Jamison Heard, Prabu David
分类: cs.CL, cs.AI, cs.CY
发布日期: 2026-08-10
备注: 9 pages, 2 figures, 2 tables. Includes technical supplement
💡 一句话要点
探讨大型语言模型如何评估社会吸引力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 社会吸引力 心理学 人类评判 性别表现 理论基础 评分稳定性
📋 核心要点
- 现有方法在评估社会吸引力时,缺乏对大型语言模型有效性的验证,尤其是在主观判断方面。
- 论文通过构建基于理论的人物档案,设计了多项研究来比较LLMs与人类的社会吸引力评估能力。
- 实验结果表明,LLMs在评分稳定性和相对排序上与人类一致,但在吸引力评分上存在显著差异。
📝 摘要(中文)
大型语言模型(LLMs)在进行传统上由人类进行的主观评估时,其作为社会评判者的有效性仍不明确。本文研究LLMs是否能够根据十个心理和关系构建的理论基础的人物档案评估社会吸引力。这些档案分为三个层次:社会吸引、社会混合和社会不吸引。通过两项研究评估LLM的评分,并在第三项研究中与人类判断进行比较。研究结果显示,尽管不同模型的评分存在差异,但在评分稳定性和相对档案排序上表现出一致性。人类参与者的评分也重现了三层结构,尽管LLMs对吸引力档案的评分普遍更高,对不吸引力档案的评分更低。性别表现对评分没有显著影响。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在社会吸引力评估中的有效性问题,现有方法缺乏对LLMs作为社会评判者的验证。
核心思路:论文通过构建基于心理和关系构建的理论人物档案,评估LLMs的评分能力,并与人类评分进行比较,以验证其有效性。
技术框架:研究分为三项主要实验:第一项是对34个LLMs进行评分,第二项是分析性别表现的敏感性,第三项是与198名人类参与者的评分进行比较。
关键创新:最重要的创新在于通过理论基础的人物档案构建,系统性地评估LLMs的社会吸引力判断能力,与传统人类评判进行对比,填补了这一领域的研究空白。
关键设计:研究中使用了12个档案,分为三类,并通过多次重复实验确保评分的稳定性,采用了性别中立的名字进行敏感性测试。
🖼️ 关键图片
📊 实验亮点
实验结果显示,LLMs在评分稳定性和相对排序上表现出一致性,且与人类评分的三层结构相符。尽管LLMs对吸引力档案的评分普遍高于人类,但两者在性别表现的影响上均未显示显著差异。
🎯 应用场景
该研究的潜在应用领域包括社交媒体内容生成、在线交互系统和人机交互设计等。通过理解LLMs如何评估社会吸引力,可以提升其在社交场景中的表现,增强用户体验,未来可能影响人机交互的设计和优化。
📄 摘要(原文)
Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.