arXiv:2502.11603cs.CLcs.AI2025-02被引 4

通过解耦性别信息与任务语义,提升大模型公平性

DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning

  • 生成无性别倾向的推理路径作为上下文示例
  • 在6个大模型上验证,核心指代消解准确率提升12.3%
  • 无需修改模型参数,适用于文本与视觉语言模型

大型语言模型虽具备强大语言理解能力,但会继承并放大社会偏见,尤其表现为性别偏见,引发公平性问题。现有基于提示的去偏方法存在关键缺陷:无法将性别信息与任务语义解耦。偏见引导会使模型过度关注性别线索,基于推理的提示则诱发性别偏见的推理链条。为此,我们提出DR.GAP(解耦推理的性别感知提示),一种自动化、模型无关的去偏流程。该方法生成性别中立的推理轨迹,并在推理时作为上下文示例使用,有效实现性别属性与任务语义的分离,且无需修改模型参数。在六个大模型上的核心指代消解和问答任务实验表明,DR.GAP在保持模型性能的同时显著降低性别偏见,具有强泛化性与鲁棒性。机制分析进一步支持其有效性。此外,该方法可扩展至视觉-语言模型(VLMs),实现显著的偏见减少。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit strong natural language understanding capabilities but also inherit and amplify societal biases, particularly gender bias, raising fairness concerns. Existing prompt-based debiasing strategies share a key limitation: they fail to disentangle gender information from task semantics. Bias steering compels models to overemphasize gender cues, while reasoning-based prompting induces gender-biased reasoning chains. To address these challenges, we propose DR.GAP (Decoupled Reasoning for Gender-Aware Prompting), an automated and model-agnostic pipeline that mitigates gender bias while preserving model performance. DR.GAP generates gender-neutral reasoning traces and applies them as in-context demonstrations during inference, effectively decoupling gender attributes from task semantics without modifying model parameters. Extensive experiments on coreference resolution and question-answering tasks across six LLMs demonstrate DR.GAP's effectiveness, generalizability, and robustness, supported by detailed mechanism analyses. Moreover, DR.GAP can be extended to vision-language models (VLMs), achieving substantial bias reduction.

大模型去偏提示工程性别公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。