平衡多模态大模型的推理能力与社会偏见,找到最佳训练比例。
FairReason: Balancing Reasoning and Social Bias in MLLMs
- 通过强化学习调整训练样本比例,实现推理与去偏的协同优化。
- 1:4的去偏与推理样本比使刻板印象降低10%,推理准确率保留88%。
- 为实际部署中的公平性与性能权衡提供可操作的实证指导。
多模态大语言模型(MLLMs)已在多种任务和模态上达到顶尖水平。为提升其推理能力,近期研究探索了高级提示策略和微调方法。尽管这些技术提高了逻辑准确性,但常导致输出中存在显著的社会偏见。厘清推理提升与偏见缓解之间的关系,以及二者是否必然存在权衡,仍是关键且紧迫的研究问题。本研究在相同条件下基准测试三种去偏策略:监督微调(SFT)、知识蒸馏(KD)和基于规则的强化学习(RL),明确其优劣。在此基础上,我们改变每种范式中去偏样本与推理样本的比例,绘制推理与偏见的权衡曲线。结果揭示出一个稳定的最佳区间:约1:4的强化学习训练比例可在降低10%刻板印象分数的同时,保留88%的原始推理准确率,为实现MLLMs的公平性与能力平衡提供了具体指导。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) already achieve state-of-the-art results across a wide range of tasks and modalities. To push their reasoning ability further, recent studies explore advanced prompting schemes and post-training fine-tuning. Although these techniques improve logical accuracy, they frequently leave the models' outputs burdened with pronounced social biases. Clarifying how reasoning gains interact with bias mitigation-and whether the two objectives inherently trade off-therefore remains an open and pressing research problem. Our study begins by benchmarking three bias-mitigation strategies-supervised fine-uning (SFT), knowledge distillation (KD), and rule-based reinforcement learning (RL)-under identical conditions, establishing their baseline strengths and weaknesses. Building on these results, we vary the proportion of debias-focused and reasoning-centric samples within each paradigm to chart the reasoning-versus-bias trade-off. Our sweeps reveal a consistent sweet spot: a roughly 1:4 mix trained with reinforcement learning cuts stereotype scores by 10% while retaining 88% of the model's original reasoning accuracy, offering concrete guidance for balancing fairness and capability in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。