Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

📄 arXiv: 2609.02512v1 📥 PDF

作者: Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, Macken Murphy

分类: cs.CV, cs.HC

发布日期: 2026-09-02


💡 一句话要点

研究表明多模态大语言模型系统性高估面部吸引力

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态大语言模型 面部吸引力 人类判断 AI模型比较 吸引力评估

📋 核心要点

  1. 现有的多模态大语言模型在面部吸引力评估中存在系统性高估的问题,无法准确反映人类的判断。
  2. 本研究通过比较人类参与者与四种商业AI模型的吸引力评分,探索MLLMs的评估机制及其与人类判断的关系。
  3. 实验结果表明,尽管MLLMs在绝对评分上与人类存在差异,但在排名顺序上表现出较强的一致性。

📝 摘要(中文)

多模态大语言模型(MLLMs)在面部吸引力评估中的应用日益普及,然而其是否能准确反映人类的吸引力判断仍然存在疑问。本文通过对2513名参与者的吸引力评分与四种商业AI模型(Claude、Gemini、GPT和Grok)的比较,发现MLLMs系统性地高估面孔的吸引力,并且评分范围较窄。尽管如此,MLLMs与人类吸引力判断之间存在较强的相关性,能够准确跟踪面孔的排名顺序。研究结果表明,当前的商业MLLMs在绝对评分上并未重现人类的评价,但在某些特征(如面部年龄)上表现出一致性。

🔬 方法详解

问题定义:本文旨在探讨多模态大语言模型在面部吸引力评估中的准确性,现有方法普遍存在高估吸引力的现象,无法与人类的判断相一致。

核心思路:通过对2513名参与者的吸引力评分与四种商业AI模型的评分进行比较,分析MLLMs的评估机制及其与人类判断的相关性。

技术框架:研究采用了预注册的探索性研究设计,收集人类参与者的评分数据,并与AI模型的评分进行系统比较,主要模块包括数据收集、评分比较和相关性分析。

关键创新:本研究的创新在于系统性地比较了不同AI模型与人类评分之间的差异,揭示了MLLMs在吸引力评估中的局限性及其潜在的评估机制。

关键设计:研究中使用了多种AI模型(Claude、Gemini、GPT、Grok),并分析了面部年龄、种族和性别等特征对吸引力评分的影响,发现仅面部年龄在两者之间表现出一致性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,MLLMs在面部吸引力评分上系统性地高估了人类的评价,且评分范围较窄。尽管如此,MLLMs与人类评分之间存在较强的相关性,能够准确反映面孔的排名顺序。Grok模型在与人类的评分一致性上表现最差。

🎯 应用场景

该研究的结果对美学、市场营销和社交媒体等领域具有重要的应用价值。理解AI模型在吸引力评估中的局限性,可以帮助相关行业更好地利用这些技术,同时也为未来的AI模型改进提供了方向。

📄 摘要(原文)

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.