用博弈论方法精准分配多路径推理奖励,提升大模型推理稳定性。
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

- 将每条推理路径视为合作博弈中的玩家,用谢尔利值量化其贡献。
- 在数学推理任务上超越基线,训练更稳定,结果更可解释。
- 适合研究大模型多步推理机制或强化学习奖励设计的学者。
大型语言模型在多步推理中表现优异,但现有并行推理方法难以区分各推理路径的实际贡献。许多路径冗余、误导甚至有害,但当前基于结果的奖励机制对所有路径统一赋值,导致学习信号模糊且训练不稳定。本文提出 Parallel Shapley,一种基于强化学习的框架,实现多路径推理中细粒度的路径级贡献分配。将每条路径视为合作博弈中的参与者,利用谢尔利值衡量边际贡献,结合生成式奖励模型评估路径效用,并通过蒙特卡洛采样实现高效近似。在数学推理基准测试中,Parallel Shapley 超越现有基线,同时带来更稳定、可解释的训练过程。该框架有效“剔除搭便车者”,按贡献比例分配奖励,显著提升大模型的多路径推理能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。