在毫米波多用户系统中,用反馈高效学习来公平分配波束和速率以满足服务质量目标。
Multi-User mmWave Beam and Rate Adaptation via Combinatorial Satisficing Bandits

- 基于满意阈值的组合半-贝叶斯方法,优先满足服务目标而非单纯最大化性能。
- 理论证明:当目标可实现时,累计满意遗憾为常数;不可实现时,瞬态期后后悔率仅随日志平方增长。
- 无需信道状态信息,实测显示在动态信道下能均衡提升吞吐量与用户公平性。
我们研究多用户毫米波MISO系统中的下行波束与速率自适应问题,多个基站(BS)使用有限码本的模拟波束成形,为多个单天线用户设备(UE)提供唯一波束及离散传输速率。基站通过ACK/NACK反馈获取传输成功信息。为编码服务目标,引入满意吞吐量阈值τ_r,将联合波束与速率自适应建模为波束-速率元组上的组合半-贝叶斯问题。在此框架下,提出SAT-CTS轻量级、阈值感知策略,融合保守置信估计与后验采样,引导学习聚焦于达成τ_r而非仅最大化性能。主要理论贡献是首次给出具有满意目标的组合半-贝叶斯问题的有限时间后悔界:当τ_r可实现时,累积满意后悔被上界控制为与时间无关的常数;当τ_r不可实现时,SAT-CTS仅在有限期望瞬态期内偏离承诺的CTS轮次,之后其后悔由重启的CTS轮次后悔之和决定,从而获得O((log T)^2)的标准后悔界。实验评估基于时变稀疏多径信道,结果表明SAT-CTS持续降低满意后悔,保持竞争性标准后悔,并在所有用户间实现良好平均吞吐量与公平性,表明无需信道状态知识的反馈高效学习可公平分配波束与速率以满足服务质量目标。
原文摘要 · Abstract (English)
We study downlink beam and rate adaptation in a multi-user mmWave MISO system where multiple base stations (BSs), each using analog beamforming from finite codebooks, serve multiple single-antenna user equipments (UEs) with a unique beam per UE and discrete data transmission rates. BSs learn about transmission success based on ACK/NACK feedback. To encode service goals, we introduce a satisficing throughput threshold $τ_r$ and cast joint beam and rate adaptation as a combinatorial semi-bandit over beam-rate tuples. Within this framework, we propose SAT-CTS, a lightweight, threshold-aware policy that blends conservative confidence estimates with posterior sampling, steering learning toward meeting $τ_r$ rather than merely maximizing. Our main theoretical contribution provides the first finite-time regret bounds for combinatorial semi-bandits with satisficing objective: when $τ_r$ is realizable, we upper bound the cumulative satisficing regret to the target with a time-independent constant, and when $τ_r$ is non-realizable, we show that SAT-CTS incurs only a finite expected transient outside committed CTS rounds, after which its regret is governed by the sum of the regret contributions of restarted CTS rounds, yielding an $O((\log T)^2)$ standard regret bound. On the practical side, we evaluate the performance via cumulative satisficing regret to $τ_r$ alongside standard regret and fairness. Experiments with time-varying sparse multipath channels show that SAT-CTS consistently reduces satisficing regret and maintains competitive standard regret, while achieving favorable average throughput and fairness across users, indicating that feedback-efficient learning can equitably allocate beams and rates to meet QoS targets without channel state knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。