Effects of Answer Format Variation on Gender Bias in Large Language Models
作者: Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber
分类: cs.CL
发布日期: 2026-08-18
备注: 6th Workshop on Computational Linguistics for the Political and Social Sciences (CPSS 2026)
💡 一句话要点
探讨回答格式变化对大型语言模型性别偏见的影响
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 性别偏见 大型语言模型 回答格式 模型评估 问卷设计
📋 核心要点
- 现有方法在评估大型语言模型的性别偏见时,未考虑回答格式对结果的影响,导致测量不准确。
- 论文通过比较不同回答格式下的模型表现,提出将回答格式作为评估的重要因素,以提高偏见测量的准确性。
- 实验结果显示,回答格式显著影响测量结果,包括排名反转,强调了多格式设计在模型评估中的重要性。
📝 摘要(中文)
大型语言模型(LLMs)中的性别偏见通常通过问答或调查基准进行评估,回答格式对结果有显著影响。本文首次研究了回答格式变化如何影响LLMs的性别偏见测量及其与人类响应分布的一致性。我们评估了三种指令调优模型在BBQ基准和OpinionQA调查数据上的表现,比较了封闭式、Likert量表和开放式格式下的偏见测量和分布一致性。结果表明,回答格式显著改变了测量结果,甚至导致排名反转,强调了将回答格式视为LLM评估的重要组成部分的必要性,并推动了多格式设计以实现更稳健的模型评估。
🔬 方法详解
问题定义:本文旨在解决回答格式变化对大型语言模型性别偏见测量的影响,现有方法未能充分考虑这一因素,导致偏见评估结果的不一致性。
核心思路:通过比较封闭式、Likert量表和开放式回答格式下的模型表现,探讨不同格式如何影响模型的响应行为和偏见测量结果。
技术框架:研究采用了三种指令调优模型,使用BBQ基准和OpinionQA调查数据进行评估,分析不同回答格式下的偏见测量和分布一致性。
关键创新:首次系统性地研究了回答格式对大型语言模型性别偏见测量的影响,揭示了不同格式下响应行为的差异及其对结果的影响。
关键设计:实验中设置了多种回答格式,采用一致的评估条件,确保比较的公平性,重点分析了各格式下的响应行为特征。
🖼️ 关键图片
📊 实验亮点
实验结果表明,回答格式对性别偏见的测量结果有显著影响,甚至导致排名反转。这一发现强调了在评估大型语言模型时,必须考虑回答格式的多样性,以实现更准确的偏见测量。
🎯 应用场景
该研究的结果对社会科学、心理学以及人工智能领域的问卷设计和模型评估具有重要意义。通过理解回答格式对偏见测量的影响,可以优化问卷设计,提高数据收集的有效性和准确性,进而推动更公平的AI系统发展。
📄 摘要(原文)
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.