arXiv:2507.23067cs.AI2025-07中稿 · ICCV被引 3

平衡多模态大模型的推理能力与社会偏见,找到最佳训练比例。

FairReason: Balancing Reasoning and Social Bias in MLLMs

  • 通过强化学习调整训练样本比例,实现推理与去偏的协同优化。
  • 1:4的去偏与推理样本比使刻板印象降低10%,推理准确率保留88%。
  • 为实际部署中的公平性与性能权衡提供可操作的实证指导。

多模态大语言模型(MLLMs)已在多种任务和模态上达到顶尖水平。为提升其推理能力,近期研究探索了高级提示策略和微调方法。尽管这些技术提高了逻辑准确性,但常导致输出中存在显著的社会偏见。厘清推理提升与偏见缓解之间的关系,以及二者是否必然存在权衡,仍是关键且紧迫的研究问题。本研究在相同条件下基准测试三种去偏策略:监督微调(SFT)、知识蒸馏(KD)和基于规则的强化学习(RL),明确其优劣。在此基础上,我们改变每种范式中去偏样本与推理样本的比例,绘制推理与偏见的权衡曲线。结果揭示出一个稳定的最佳区间:约1:4的强化学习训练比例可在降低10%刻板印象分数的同时,保留88%的原始推理准确率,为实现MLLMs的公平性与能力平衡提供了具体指导。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) already achieve state-of-the-art results across a wide range of tasks and modalities. To push their reasoning ability further, recent studies explore advanced prompting schemes and post-training fine-tuning. Although these techniques improve logical accuracy, they frequently leave the models' outputs burdened with pronounced social biases. Clarifying how reasoning gains interact with bias mitigation-and whether the two objectives inherently trade off-therefore remains an open and pressing research problem. Our study begins by benchmarking three bias-mitigation strategies-supervised fine-uning (SFT), knowledge distillation (KD), and rule-based reinforcement learning (RL)-under identical conditions, establishing their baseline strengths and weaknesses. Building on these results, we vary the proportion of debias-focused and reasoning-centric samples within each paradigm to chart the reasoning-versus-bias trade-off. Our sweeps reveal a consistent sweet spot: a roughly 1:4 mix trained with reinforcement learning cuts stereotype scores by 10% while retaining 88% of the model's original reasoning accuracy, offering concrete guidance for balancing fairness and capability in MLLMs.

多模态去偏强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。