提出递归梯度机制,让智能体在合作博弈中高效计算声誉影响。
The Reciprocity Gradient

- 通过解析传播声誉链,反向传播奖励梯度
- 在无奖励设计下恢复近最优的上下文敏感策略
- 适合研究合作博弈与声誉建模的研究者
沟通是维系战略互动中互惠与合作的基础。我们识别并形式化了学习智能体在动态中的核心优化难题:智能体发出的任何行动或信号,都会通过组合式分支路径重塑多个第三方的声誉,并反馈至自身未来的奖励,迫使智能体在每一步决策时必须同时考虑所有这些间接通道。为此,我们提出递归梯度(Reciprocity Gradient),该方法通过从公开观测中训练的对手策略私有估计器,显式地将奖励梯度反向传播通过声誉链本身,而非依赖采样回报估算。该梯度分析性地流经声誉链,联合优化动作与评估信号,无需内在奖励或奖励塑形。实验表明,该方法能恢复近最优的上下文敏感策略,而基于样本的基线方法则退化为恒定输出策略。
原文摘要 · Abstract (English)
Communication is fundamental to sustaining reciprocity and cooperation in strategic interactions. We identify and formulate the influence attribution problem as the central optimization difficulty inherent in such dynamics for a learning agent: any action or signal the agent emits reshapes the reputations of many third parties along combinatorially branching paths before feeding back into its own future rewards, forcing the agent to account for all of these indirect channels at once when choosing every action. To address this, we introduce the reciprocity gradient, which explicitly backpropagates reward gradients through private estimators of opponents' policies trained from public observations. The gradient flows through the reputation chain itself analytically, rather than being estimated from sampled returns. It jointly optimizes actions and evaluative signals without intrinsic rewards or reward shaping. Empirically, the method recovers near-optimal context-sensitive policies, while sample-based baselines collapse into constant-output policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。