arXiv:2412.11414cs.CLcs.LG2024-12被引 4

通过修复理解错误,让大模型少说刻板印象

Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws

  • 用指令微调修复理解偏差,间接减少刻板输出
  • 多维度刻板印象减少超60%,无需显式去偏技术
  • 适合关注模型真实偏差与错误区分的研究者

近期研究表明,生成式语言模型常反映并放大社会偏见。但此类研究常将偏见与任务特定缺陷(如理解失败)混淆。例如,当模型误解文本并产生强化刻板印象的回复时,难以判断问题源于内在偏见还是内容理解错误。本文通过多维度评估,明确区分阅读理解任务中的偏见与缺陷。提出一种针对性的刻板印象缓解框架,通过在通用数据集上进行指令微调,隐式降低生成模型中的刻板印象。该方法在国籍、年龄、性别、残疾及外貌等多个维度上,使刻板输出减少超过60%,且不依赖显式去偏技术。我们在多个主流生成模型上验证了该方法的有效性,同时保持整体性能。研究强调必须严格区分‘偏见’与其他错误,才能制定更精准的缓解策略。内容警告:部分示例包含冒犯性刻板印象。

原文摘要 · Abstract (English)

Recent studies have shown that generative language models often reflect and amplify societal biases in their outputs. However, these studies frequently conflate observed biases with other task-specific shortcomings, such as comprehension failure. For example, when a model misinterprets a text and produces a response that reinforces a stereotype, it becomes difficult to determine whether the issue arises from inherent bias or from a misunderstanding of the given content. In this paper, we conduct a multi-faceted evaluation that distinctly disentangles bias from flaws within the reading comprehension task. We propose a targeted stereotype mitigation framework that implicitly mitigates observed stereotypes in generative models through instruction-tuning on general-purpose datasets. We reduce stereotypical outputs by over 60% across multiple dimensions -- including nationality, age, gender, disability, and physical appearance -- by addressing comprehension-based failures, and without relying on explicit debiasing techniques. We evaluate several state-of-the-art generative models to demonstrate the effectiveness of our approach while maintaining the overall utility. Our findings highlight the need to critically disentangle the concept of `bias' from other types of errors to build more targeted and effective mitigation strategies. CONTENT WARNING: Some examples contain offensive stereotypes.

刻板印象指令微调模型评估去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。