arXiv:2502.11559cs.CLcs.AI2025-02NeurIPS被引 3

自动生成公平指令,让大模型回答更公正

Auto-Search and Refinement: An Automated Framework for Gender Bias Mitigation in Large Language Models

  • 用自动搜索与迭代优化生成公平提示词
  • 在不改模型的前提下显著降低性别偏见
  • 适配闭源和开源模型,兼顾公平与性能

在海量文本上预训练的大语言模型虽提升自然语言处理能力,但可能内化社会偏见,尤其是性别偏见。现有参数调整方法如微调成本高,不适用于闭源模型,且难以适应社会规范变化;基于指令的方法虽灵活,但常损害任务表现。为此,我们提出 FaIRMaker,一种无需依赖模型的自动化框架,采用自动搜索与迭代优化机制,动态生成名为 Fairwords 的公平指令,嵌入输入以缓解性别偏见并提升输出质量。大量实验表明,该框架可自动搜索并持续优化公平指令,在保持任务完整性的同时有效减少性别偏见,兼容基于 API 和开源的大模型。

原文摘要 · Abstract (English)

Pre-training large language models (LLMs) on vast text corpora enhances natural language processing capabilities but risks encoding social biases, particularly gender bias. While parameter-modification methods like fine-tuning mitigate bias, they are resource-intensive, unsuitable for closed-source models, and lack adaptability to evolving societal norms. Instruction-based approaches offer flexibility but often compromise task performance. To address these limitations, we propose $\textbf{FaIRMaker}$, an automated and model-independent framework that employs an $\textbf{auto-search and refinement}$ paradigm to adaptively generate Fairwords, which act as instructions integrated into input queries to reduce gender bias and enhance response quality. Extensive experiments demonstrate that FaIRMaker automatically searches for and dynamically refines Fairwords, effectively mitigating gender bias while preserving task integrity and ensuring compatibility with both API-based and open-source LLMs.

性别偏见指令优化自动化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。