研究大模型在回答性别相关问题时的信息量差异,发现个体问题上存在偏差但整体被抵消。
Do LLMs have a Gender (Entropy) Bias?
- 通过真实世界问题测试大模型对男女提问的信息输出差异
- 78%情况下,去偏策略生成的内容信息量更高
- 适合关注AI公平性与提示工程的研究者和开发者
我们研究了部分主流大模型中是否存在一种特定的性别偏差,并构建了一个新的基准数据集RealWorldQuestioning(发布于HuggingFace),该数据集来自商业与健康领域中的教育、就业、个人财务管理及一般健康四个关键方向的真实问题。我们定义并分析了熵偏差(entropy bias),即大模型在回应不同性别用户实际提问时所生成信息量的差异。通过对四种大模型进行测试,并使用ChatGPT-4o作为“模型裁判”进行定性和定量评估,结果显示,在类别层面,男性与女性的响应无显著差异;但在更细粒度的个体问题层面,绝大多数情况下存在明显差异,且常因部分响应对男性更有利、部分对女性更有利而相互抵消。这对普通用户而言仍是隐患,因用户通常只提一个具体问题。为此我们提出一种简单的迭代合并双性别响应的去偏方法,该策略在78%案例中生成的信息量超过任一性别原始响应,其余情况也实现了稳定平衡整合。
原文摘要 · Abstract (English)
We investigate the existence and persistence of a specific type of gender bias in some of the popular LLMs and contribute a new benchmark dataset, RealWorldQuestioning (released on HuggingFace ), developed from real-world questions across four key domains in business and health contexts: education, jobs, personal financial management, and general health. We define and study entropy bias, which we define as a discrepancy in the amount of information generated by an LLM in response to real questions users have asked. We tested this using four different LLMs and evaluated the generated responses both qualitatively and quantitatively by using ChatGPT-4o (as "LLM-as-judge"). Our analyses (metric-based comparisons and "LLM-as-judge" evaluation) suggest that there is no significant bias in LLM responses for men and women at a category level. However, at a finer granularity (the individual question level), there are substantial differences in LLM responses for men and women in the majority of cases, which "cancel" each other out often due to some responses being better for males and vice versa. This is still a concern since typical users of these tools often ask a specific question (only) as opposed to several varied ones in each of these common yet important areas of life. We suggest a simple debiasing approach that iteratively merges the responses for the two genders to produce a final result. Our approach demonstrates that a simple, prompt-based debiasing strategy can effectively debias LLM outputs, thus producing responses with higher information content than both gendered variants in 78% of the cases, and consistently achieving a balanced integration in the remaining cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。