arXiv:2412.01711cs.CL2024-12被引 1

用小模型生成去偏信号,高效可解释地降低大模型偏见。

Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models

  • 用小型偏见与反偏见模型生成去偏信号,在解码时修正大模型输出。
  • 在多个架构上减少性别、种族和宗教偏见,同时保持语言模型性能。
  • 方法高效可解释,适合需要可控去偏的场景,如医疗或司法应用。

尽管大型语言模型在众多应用中表现优异,但其仍会继承训练数据中的偏见,可能对边缘群体造成伤害。本文通过微调小型偏见模型和反偏见模型,生成解码时添加的去偏信号,实现高效且可解释的偏见缓解。该方法避免了对大模型重新训练,计算成本更低;同时可分析去偏过程中的概率变化,提升可解释性。通过更换微调数据集,该框架可适配特定应用场景。在多种模型架构上,对性别、种族和宗教偏见的实验表明,该方法在多个局部和全局偏见指标上均有效降低偏见,同时保持语言模型性能。

原文摘要 · Abstract (English)

Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in the training data, potentially leading to harm for marginalized communities. In this paper, we mitigate bias by leveraging small biased and anti-biased expert models to obtain a debiasing signal that is added to the LLM output at decoding-time. This approach combines computational efficiency - fine-tuning a small model versus re-training a large model and interpretability - one can examine the probability shift from debiasing. The framework can also be tailored to specific contexts by switching the choice of the fine-tuning dataset. Experiments on mitigating gender, race, and religion biases on different architectures show a reduction in bias on several local and global bias metrics while preserving language model performance.

去偏大模型可解释高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。