用公平性评分模型引导大模型推理,提升决策公正性。
Guiding LLM Decision-Making with Fairness Reward Models
- 训练公平性评分模型,评估大模型推理过程的公正性
- 跨任务跨领域无需微调,显著提升决策公平性
- 适合需公正决策的高风险场景如贷款、量刑
大型语言模型越来越多地用于支持高风险决策,可能影响谁获得保释或贷款。简单的思维链采样虽能提高平均决策准确率,但会加剧不公平偏见。为此,我们提出一种通用公平性奖励模型(FRM)训练框架。该模型为大模型的推理过程打分,使系统在聚合多个推理路径时降低偏见路径权重,倾向更公平的结果。我们证明,一个在弱监督、由大模型标注的偏见与无偏见推理样本上训练的单一公平性奖励模型,可在不额外微调的情况下跨任务、跨领域和跨模型家族迁移。应用于实际决策任务如再犯预测和社会媒体内容审核,我们的方法在保持或超越基线准确率的同时,持续提升公平性。
原文摘要 · Abstract (English)
Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to amplify unfair bias. To address this challenge and enable the trustworthy use of reasoning models in high-stakes decision-making, we propose a framework for training a generalizable Fairness Reward Model (FRM). Our model assigns a fairness score to LLM reasoning, enabling the system to down-weight biased trajectories and favor equitable ones when aggregating decisions across reasoning chains. We show that a single Fairness Reward Model, trained on weakly supervised, LLM-annotated examples of biased versus unbiased reasoning, transfers across tasks, domains, and model families without additional fine-tuning. Applied to real-world decision-making tasks including recidivism prediction and social media moderation, we show that our approach consistently improves fairness while matching, or even surpassing, baseline accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。