arXiv:2502.15240cs.LGcs.SY2025-02

兼顾公平与效率的多智能体抽奖算法,可保障最低收益

Multi-agent Multi-armed Bandits with Minimum Reward Guarantee Fairness

  • 基于上置信界改进,动态平衡收益与公平性
  • 社会福利与公平性误差均低于√T量级,理论最优
  • 适合需公平分配资源的推荐系统与调度场景

我们研究在多智能体多臂老虎机(MA-MAB)设置中,在保障公平性的前提下最大化社会福利的问题。中心决策者随时间采取行动,为不同智能体产生随机回报。目标是最大化期望累积回报之和(即社会福利),同时确保每个智能体的期望回报不低于最大可能期望回报的固定比例。提出的RewardFairUCB算法利用上置信界(UCB)技术,实现了社会福利与公平性两方面的次线性后悔边界。公平性后悔衡量的是最小奖励保障与实际政策回报之间的正差值,社会福利后悔则衡量最优公平策略与当前策略之间的差距。我们证明RewardFairUCB实现社会福利后悔的实例无关上界为˜O(T^{1/2}),公平性后悔上界为˜O(T^{3/4})。同时给出社会福利与公平性后悔的Ω(√T)下界。通过模拟数据与真实数据对比多种基线算法,揭示了公平性与社会福利之间的权衡。

原文摘要 · Abstract (English)

We investigate the problem of maximizing social welfare while ensuring fairness in a multi-agent multi-armed bandit (MA-MAB) setting. In this problem, a centralized decision-maker takes actions over time, generating random rewards for various agents. Our goal is to maximize the sum of expected cumulative rewards, a.k.a. social welfare, while ensuring that each agent receives an expected reward that is at least a constant fraction of the maximum possible expected reward. Our proposed algorithm, RewardFairUCB, leverages the Upper Confidence Bound (UCB) technique to achieve sublinear regret bounds for both fairness and social welfare. The fairness regret measures the positive difference between the minimum reward guarantee and the expected reward of a given policy, whereas the social welfare regret measures the difference between the social welfare of the optimal fair policy and that of the given policy. We show that RewardFairUCB algorithm achieves instance-independent social welfare regret guarantees of $\tilde{O}(T^{1/2})$ and a fairness regret upper bound of $\tilde{O}(T^{3/4})$. We also give the lower bound of $Ω(\sqrt{T})$ for both social welfare and fairness regret. We evaluate RewardFairUCB's performance against various baseline and heuristic algorithms using simulated data and real world data, highlighting trade-offs between fairness and social welfare regrets.

多智能体公平性强化学习优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。