通过引导推理提升语言模型公平性,有效减少刻板印象响应。
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
- 用高级模型的推理路径指导低能力模型,无需专门标注公平性数据。
- 改进推理后,模型在公平性评测中表现超越原先进阶模型。
- 推理正确性和长度直接影响模型公平性与整体性能。
近期大模型进展表明,推理能力能显著提升模型在多种任务上的表现。然而,推理对缓解刻板偏见的影响仍不明确。本文研究模型推理能力与公平性的关系,发现具备更强推理能力的大模型在现有公平性基准上表现出显著更低的刻板偏见。基于此,提出ReGiFT(Reasoning Guided Fine-Tuning)方法:从先进推理模型中提取结构化推理轨迹,并注入缺乏此类能力的模型中。该方法仅依赖通用推理能力,无需任何特定于公平性的监督信号。实验显示,经ReGiFT微调的模型不仅比非推理型模型更公平,且在公平性评测中优于部分先进推理模型。我们还分析了推理轨迹的正确性与长度对模型公平性及整体性能的影响。结果表明,增强推理能力是一种有效的、不依赖公平性标注的偏见缓解策略,尤其针对由浅层或错误推理引发的刻板印象。
原文摘要 · Abstract (English)
Recent advances in large-scale generative language models have shown that reasoning capabilities can significantly improve model performance across a variety of tasks. However, the impact of reasoning on a model's ability to mitigate stereotypical responses remains largely underexplored. In this work, we investigate the crucial relationship between a model's reasoning ability and fairness, and ask whether improved reasoning capabilities can mitigate harmful stereotypical responses, especially those arising due to shallow or flawed reasoning. We conduct a comprehensive evaluation of multiple open-source LLMs, and find that larger models with stronger reasoning abilities exhibit substantially lower stereotypical bias on existing fairness benchmarks. Building on this insight, we introduce ReGiFT -- Reasoning Guided Fine-Tuning, a novel approach that extracts structured reasoning traces from advanced reasoning models and infuses them into models that lack such capabilities. We use only general-purpose reasoning and do not require any fairness-specific supervision for bias mitigation. Notably, we see that models fine-tuned using ReGiFT not only improve fairness relative to their non-reasoning counterparts but also outperform advanced reasoning models on fairness benchmarks. We also analyze how variations in the correctness of the reasoning traces and their length influence model fairness and their overall performance. Our findings highlight that enhancing reasoning capabilities is an effective, fairness-agnostic strategy for mitigating stereotypical bias caused by reasoning flaws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。