Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study
作者: Simone Mungari
分类: cs.CL
发布日期: 2026-08-12
💡 一句话要点
提出审计框架以评估大型语言模型的政治倾向
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 政治倾向 审计框架 意大利案例 模型评估 政治信息 用户影响
📋 核心要点
- 现有研究未能系统性地审计大型语言模型在政治问题上的表现,导致对其影响的理解不足。
- 本文提出了一种审计框架,允许对多个LLMs在评估政党和领导人时的行为进行系统分析。
- 通过意大利案例研究,展示了LLMs在不同提示下的评估一致性和差异性,提供了重要的实证数据。
📝 摘要(中文)
随着用户在政治事务中越来越依赖大型语言模型(LLMs)获取信息和建议,模型所表达的政治偏好引发了公众关注。本文研究了LLMs如何表达对政党和政治领导人的偏好,提出了一种系统且可重复的审计框架,促使多个LLMs在九个标准下评估政党和领导人。我们关注模型的可观察行为,分析评估的一致性、模型间的差异、拒绝率及对提示形式的敏感性,并探讨在不同角色指令下的评估变化。通过意大利案例研究,系统分析了LLM生成的对意大利政党和领导人的政治评估。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在政治评估中的偏见和不一致性问题。现有方法缺乏系统性审计,无法有效评估模型的政治倾向。
核心思路:论文提出了一种系统的审计框架,通过对多个LLMs进行统一的评估,分析其在不同提示下的表现,关注可观察的行为而非模型的“真实”信念。
技术框架:整体框架包括数据收集、模型选择、评估标准设定和结果分析四个主要模块。每个模块都有明确的目标和方法,以确保审计过程的系统性和可重复性。
关键创新:最重要的创新在于提出了一种可重复的审计方法,能够系统性地比较不同LLMs的政治评估,填补了现有研究的空白。
关键设计:在评估过程中,设置了九个标准用于比较,关注模型的评估一致性、拒绝率和对提示的敏感性等关键参数。
🖼️ 关键图片
📊 实验亮点
实验结果表明,不同LLMs在政治评估上的一致性存在显著差异,某些模型在特定提示下的拒绝率高达30%。通过系统分析,本文揭示了模型对不同政治角色的敏感性,为理解LLMs的政治倾向提供了实证依据。
🎯 应用场景
该研究的潜在应用领域包括政治咨询、舆情分析和社交媒体监测等。通过理解LLMs的政治倾向,可以帮助用户更好地识别和应对信息偏见,从而提高公众对政治信息的判断能力。未来,该框架可扩展至其他领域的模型审计,提升模型透明度与可信度。
📄 摘要(原文)
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors. In this paper, we investigate whether and how LLMs express preferences toward political parties and political leaders. We introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria. Rather than attempting to infer the models' "true" political beliefs, we focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation. We further investigate how these evaluations vary when models are instructed to adopt different personas. We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.