arXiv:2410.17519cs.CL2024-10ACL被引 18

发现大模型在长文本生成中存在隐性偏见,提出新评测框架并改进方法。

Large Language Models Still Exhibit Bias in Long Text

  • 设计长文本公平性测试框架,覆盖14主题10类人群
  • 发现模型对特定群体偏好且过度保护弱势群体
  • 提出微调方法,降低34.6%性别偏见,提升评测表现

现有大模型公平性评测多集中于选择题等简单任务,忽视长文本生成中的潜在偏见。为此,我们提出长文本公平性测试(LTF-TEST)框架,通过议论文式提示评估模型偏见,涵盖14个主题和10个人口学维度(包括性别、种族),生成11,948个样本。该框架不仅分析模型输出,还考察其推理过程,揭示了传统简单响应难以捕捉的细微偏见。在评估GPT-4o、LLaMa3等五款近期大模型时,发现两大偏见模式:一是模型频繁在回应中倾向特定群体;二是对传统弱势群体表现出过度敏感,常给出过度保护性回应而忽略其他群体。为缓解此问题,我们提出FT-REGARD微调方法,将有偏提示与中立回应配对。该方法使性别偏见降低34.6%,并在BBQ基准上性能提升1.4个百分点,为解决长文本生成中的偏见问题提供了有效路径。

原文摘要 · Abstract (English)

Existing fairness benchmarks for large language models (LLMs) primarily focus on simple tasks, such as multiple-choice questions, overlooking biases that may arise in more complex scenarios like long-text generation. To address this gap, we introduce the Long Text Fairness Test (LTF-TEST), a framework that evaluates biases in LLMs through essay-style prompts. LTF-TEST covers 14 topics and 10 demographic axes, including gender and race, resulting in 11,948 samples. By assessing both model responses and the reasoning behind them, LTF-TEST uncovers subtle biases that are difficult to detect in simple responses. In our evaluation of five recent LLMs, including GPT-4o and LLaMa3, we identify two key patterns of bias. First, these models frequently favor certain demographic groups in their responses. Second, they show excessive sensitivity toward traditionally disadvantaged groups, often providing overly protective responses while neglecting others. To mitigate these biases, we propose FT-REGARD, a finetuning approach that pairs biased prompts with neutral responses. FT-REGARD reduces gender bias by 34.6% and improves performance by 1.4 percentage points on the BBQ benchmark, offering a promising approach to addressing biases in long-text generation tasks.

大模型偏见长文本生成公平性评测微调方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。