leader通过调整收益分配,动态学习并影响跟随者的多目标偏好。
Learning in Repeated Multi-Objective Stackelberg Games with Payoff Manipulation
- 设计基于期望收益的策略,平衡短期获利与长期偏好获取
- 在无限重复博弈中,长期期望收益策略收敛至最优操纵方案
- 无需谈判或先验知识,可提升双方共赢结果
我们研究重复多目标斯塔克尔伯格博弈中的收益操纵问题,其中领导者可通过向追随者提供自身收益份额等方式,战略性地影响其确定性最优响应。假设追随者的效用函数为线性但未知,其目标权重需通过交互推断。这给领导者带来序列决策挑战:需在偏好获取与即时收益间权衡。本文形式化该问题,提出基于期望效用(EU)与长期期望效用(longEU)的操纵策略,指导领导者选择行动与激励方案,以权衡短期收益与长期影响。理论证明,在无限重复交互下,longEU策略收敛至最优操纵。基准环境下的实证结果表明,该方法在不依赖显式协商或先验知识的前提下,显著提升领导者累计效用,并促进互利结果。
原文摘要 · Abstract (English)
We study payoff manipulation in repeated multi-objective Stackelberg games, where a leader may strategically influence a follower's deterministic best response, e.g., by offering a share of their own payoff. We assume that the follower's utility function, representing preferences over multiple objectives, is unknown but linear, and its weight parameter must be inferred through interaction. This introduces a sequential decision-making challenge for the leader, who must balance preference elicitation with immediate utility maximisation. We formalise this problem and propose manipulation policies based on expected utility (EU) and long-term expected utility (longEU), which guide the leader in selecting actions and offering incentives that trade off short-term gains with long-term impact. We prove that under infinite repeated interactions, longEU converges to the optimal manipulation. Empirical results across benchmark environments demonstrate that our approach improves cumulative leader utility while promoting mutually beneficial outcomes, all without requiring explicit negotiation or prior knowledge of the follower's utility function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。