为多大模型协作系统设计可审计的奖惩信号,让整体评价精准转化为个体行为指导。
Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- 融合博弈论归因与过程奖励建模,生成局部、有符号且信用守恒的训练信号
- 成功时公平分配贡献,失败时定位首错并鼓励修正,信号可直接用于强化或偏好训练
- 提供从全局评估到个体行为的统一监督路径,适合需可解释协作的多智能体场景
多大语言模型协作系统在处理复杂任务上展现出潜力,但现有训练方法缺乏将系统级评估与代理及消息级学习相连接的合理机制。本文提出一个理论框架,将合作博弈论归因与过程奖励建模相结合,实现从系统评价到代理与响应级信号的转化。与仅依赖归因(如Shapley值)或步骤级标签(如PRM)的方法不同,本方法生成局部、有符号且信用守恒的信号。在成功案例中,基于Shapley的信用分配公平地分配成果,并细化为每条消息的奖励,促进合作同时抑制冗余或破坏行为;在失败案例中,首次错误定位产生修复感知偏好,惩罚有害步骤并奖励纠正尝试。生成的信号具有界性、合作性,且可直接兼容强化学习或偏好基后训练,为大模型多智能体训练提供了统一且可审计的从全局评估到局部监督的路径。本文贡献在于概念层面:提出理论基础与训练信号,实证验证留待未来工作。
原文摘要 · Abstract (English)
Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principled ways to connect system-level evaluation with agent- and message-level learning. We propose a theoretical framework that unifies cooperative game-theoretic attribution with process reward modeling to transform system evaluation to agent credit to response-level signals. Unlike prior approaches that rely only on attribution (Shapley) or step-level labels (PRM), our method produces local, signed, and credit-conserving signals. In success cases, Shapley-based credit assignment fairly allocates outcomes across agents and is refined into per-message rewards that promote cooperation while discouraging redundancy or sabotage; in failure cases, first-error localization yields repair-aware preferences that penalize harmful steps while rewarding corrective attempts. The resulting signals are bounded, cooperative, and directly compatible with reinforcement- or preference-based post-training, providing a unified and auditable pathway from global evaluation to local supervision in LLM multi-agent training. Our contribution is conceptual: we present a theoretical foundation and training signals, leaving empirical validation for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。