测试方式影响大模型性别偏见测量结果,提示评估设计需更谨慎。
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
- 通过改变提示语的评估显性程度,观察模型输出变化。
- 显性评估提示使性别输出差异增大,离散选择比概率更放大偏见。
- 研究提醒评估基准需考虑生态效度,适合评测方法开发者参考。
随着大模型在社会关键场景中的应用日益广泛,性别偏见问题引发广泛关注,相关研究多聚焦于测量与缓解此类偏见。现有评估常依赖与自然语言分布不同的任务设计,通常使用明确或隐含引导性别偏见内容的提示。本文研究评估任务的提示信号如何影响大模型性别偏见的测量结果。具体地,在四种任务格式下,我们对比了(1)凸显测试情境、(2)突出性别相关内容的提示条件,并采用词元概率与离散选择两种度量方式评估模型敏感性。结果表明,越接近性别偏见评估框架的提示,越引发显著不同的性别输出分布;离散选择指标相比概率度量进一步放大偏见。这些发现不仅揭示了大模型性别偏见评估的脆弱性,也向自然语言处理评测与开发社区提出新问题:精心设计的测试能否触发模型‘应试模式’?这对未来评测基准的生态效度意味着什么?
原文摘要 · Abstract (English)
As LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias. These efforts often rely on evaluation tasks that differ from natural language distributions, as they typically involve carefully constructed task prompts that overtly or covertly signal the presence of gender bias-related content. In this paper, we examine how signaling the evaluative purpose of a task impacts measured gender bias in LLMs. Concretely, we test models under prompt conditions that (1) make the testing context salient, and (2) make gender-focused content salient. We then assess prompt sensitivity across four task formats with both token-probability and discrete-choice metrics. We find that prompts that more clearly align with (gender bias) evaluation framing elicit distinct gender output distributions compared to less evaluation-framed prompts. Discrete-choice metrics further tend to amplify bias relative to probabilistic measures. These findings do not only highlight the brittleness of LLM gender bias evaluations but open a new puzzle for the NLP benchmarking and development community: To what extent can well-controlled testing designs trigger LLM "testing mode" performance, and what does this mean for the ecological validity of future benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。