arXiv:2608.22992cs.NIcs.LG2026-08

TACAN用通道令牌注意力提升突发主用户场景下的频谱接入可靠性。

Channel-Token Attention for Reliable Dynamic Spectrum Access under Bursty Primary-User Traffic

  • 将每个信道建模为含占用历史和调制熵的令牌,结合上下文信息进行动态决策。
  • 在20信道网络中实现92.53%的包成功接入率,优于贪心策略2.59个百分点。
  • 特别在极端负载下性能提升显著,且降低平均延迟与可靠性差距。

动态频谱接入需在突发主用户活动下协调次级用户,同时保障包可靠性和时延。本文提出TACAN,一种集中式策略,将每个信道表示为包含占用历史和自动调制分类熵的令牌,上下文令牌则包含队列类别、时延和用户身份。使用基于占用贪婪策略预热的Transformer编码器,并通过近端策略优化进行精炼。冻结策略在每时隙均保持信道分配,包括队列为空时。因此在保留轨迹上回放,区分待机分配成功与有包接入及交付。在20信道网络中,60个主设备与4个次级用户下,TACAN实现92.53% ± 0.47的包存在接入成功率,优于贪心策略(89.94%)和PPO+MLP(83.53%)。相对于贪心策略的提升为2.59点(95%置信区间1.89–3.29),所有五种子实验均获胜;双边符号检验值为0.0625。该增益在正常负载下为0.57点,极端负载下升至7.67点。同时,平均交付时延从1.208降至1.123时隙,条件用户可靠性差距由9.69降至3.15点。每位次级用户每时隙传送包数仍受限于到达率(30.12%对比30.11%),故不宣称吞吐量提升。

原文摘要 · Abstract (English)

Dynamic spectrum access must coordinate secondary users under bursty primary-user activity while preserving packet reliability and delay. We present TACAN, a centralized policy that represents each channel as a token containing occupancy history and automatic-modulation-classification entropy; a context token supplies queue class, delay and user identity. A Transformer encoder is warm-started from an occupancy-greedy policy and refined with proximal policy optimization. The frozen policies were trained to maintain a channel assignment in every slot, including when queues were empty. We therefore replay them on held-out trajectories and distinguish standby assignment success from packet-present access and packet delivery. In a 20-channel network with 60 primary devices and 4 secondary users, TACAN achieves 92.53% +/- 0.47 packet-present access success, compared with 89.94% for Greedy and 83.53% for PPO+MLP. Its paired gain over Greedy is 2.59 points (parametric 95% CI 1.89-3.29), with wins in all five seeds; the exact two-sided sign-test value is 0.0625. The gain rises from 0.57 points at normal primary-user load to 7.67 points at extreme load. TACAN also reduces mean delivery delay from 1.208 to 1.123 slots and the conditional user-reliability gap from 9.69 to 3.15 points. Delivered packets per SU-slot remain arrival-limited (30.12% versus 30.11%), so no packet-throughput gain is claimed.

频谱接入注意力机制强化学习动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。