arXiv:2608.17516cs.CL2026-08

答案格式变化会显著影响大模型性别偏见的测量结果。

Effects of Answer Format Variation on Gender Bias in Large Language Models

论文配图:Effects of Answer Format Variation on Gender Bias in Large Language Models
图 1 · 摘自论文原文
  • 对比封闭、量表和开放三种答案格式,考察其对偏见测量的影响。
  • 不同格式导致偏见排序反转,如男性更常被选为领导。
  • 提醒评估时需多格式设计,避免结果偏差,适合模型评测研究者。

大语言模型(LLMs)的性别偏见常通过问答或调查基准进行评估,要求模型以预设格式作答。在调查科学中,答案格式对回应有显著影响,而LLMs也对提示词敏感。然而,目前尚不清楚答案格式如何影响性别偏见的测量及其与人类响应分布的一致性。我们使用三个指令微调模型,在BBQ基准和OpinionQA调查数据上,比较封闭式、李克特量表和开放式三种格式下的偏见测量结果与分布对齐情况。在其他条件相同的情况下,发现答案格式显著改变测量结果,甚至导致偏见排名反转。这是因为不同格式引发不同的响应行为:强制选择、量表分布和自由文本拒绝。研究强调应将答案格式视为评估的重要组成部分,建议采用多格式设计以实现更稳健的模型评估。

原文摘要 · Abstract (English)

Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.

偏见评估大模型答案格式评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。