arXiv:2509.12457cs.LG2025-09

提出兼顾公平与稳定性的组合多臂赌博机算法,提升无线网络资源调度效率。

On the Regularity and Fairness of Combinatorial Multi-Armed Bandit

  • 融合队列长度、奖励间隔与置信上界,动态调整选臂策略
  • 理论证明公平性、规律性与累积后悔量均为零
  • 适合需保障资源分配公平与稳定性的无线网络场景

组合多臂赌博机模型旨在通过每轮激活一组臂来最大化累积奖励。本文受无线网络中两类关键应用启发:既要最大化累积奖励,又要保证各臂的最小平均奖励(公平性)以及奖励发放的规律性(即每臂多久获得一次奖励)。为此,本文提出一种参数化正则且公平的学习算法,其权重度量线性结合了虚拟队列长度(追踪公平性偏差)、时间-上次奖励(TSLR)指标(衡量自上次奖励以来的轮数,反映奖励规律性)和置信上界(UCB)估计(平衡探索与利用)。通过揭示虚拟队列长度与TSLR之间的关键关系,并设计多个非平凡的李雅普诺夫函数,本文在理论上严格证明了所提算法可实现零累积公平性违规、奖励规律性及累积后悔性能。仿真结果基于两个真实数据集验证了上述理论结论。

原文摘要 · Abstract (English)

The combinatorial multi-armed bandit model is designed to maximize cumulative rewards in the presence of uncertainty by activating a subset of arms in each round. This paper is inspired by two critical applications in wireless networks, where it's not only essential to maximize cumulative rewards but also to guarantee fairness among arms (i.e., the minimum average reward required by each arm) and ensure reward regularity (i.e., how often each arm receives the reward). In this paper, we propose a parameterized regular and fair learning algorithm to achieve these three objectives. In particular, the proposed algorithm linearly combines virtual queue-lengths (tracking the fairness violations), Time-Since-Last-Reward (TSLR) metrics, and Upper Confidence Bound (UCB) estimates in its weight measure. Here, TSLR is similar to age-of-information and measures the elapsed number of rounds since the last time an arm received a reward, capturing the reward regularity performance, and UCB estimates are utilized to balance the tradeoff between exploration and exploitation in online learning. By exploring a key relationship between virtual queue-lengths and TSLR metrics and utilizing several non-trivial Lyapunov functions, we analytically characterize zero cumulative fairness violation, reward regularity, and cumulative regret performance under our proposed algorithm. These theoretical outcomes are verified by simulations based on two real-world datasets.

多臂赌博机公平性无线网络在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。