用提示词驱动的决策变换器,让无线网络多任务学习更高效泛化。
Generalizable Multi-Task Learning for Wireless Networks Using Prompt Decision Transformers

- 将多小区选择转化为序列建模问题,利用提示词实现跨场景学习。
- 在多任务设置下,用户体验提升最高达49%,且模型越大效果越好。
- 无需重训练即可快速适应新网络配置,适合动态变化的未来无线网络。
未来无线网络需快速适应高度异构环境与动态任务配置,推动无线资源管理(RRM)从传统规则和优化方法转向人工智能驱动模式。AI方法可学习复杂非线性关系,在多样化网络条件下泛化,并实现实时、可扩展、自主决策。其中,协同多点(CoMP)传输对缓解小区间干扰、提升边缘用户性能至关重要。然而,最优多小区选择是复杂的组合难题,需在动态流量与信道条件下联合优化大量可能的服务小区组合。尽管深度强化学习(DRL)如近端策略优化(PPO)取得成功,仍存在样本效率低、泛化能力差、状态与动作空间变化时需昂贵重训练等问题。为此,本文提出基于提示词决策变换器(PromptDT)的多任务学习框架,可在多种网络配置下学习,并将多小区选择重构为序列建模问题。通过利用离线轨迹与任务特定提示,PromptDT实现了对不同基站数量、终端设备数及调度策略的可扩展学习。实验表明,在多任务设置下,PromptDT相较基线提升最高达49%的用户体验(QoE),性能随模型容量正向增长。此外,其能有效泛化至未见任务,在无需重训练或微调的情况下实现稳健的少样本适应。
原文摘要 · Abstract (English)
Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio resource management (RRM) toward artificial intelligence (AI)-driven RRM. AI-driven approaches can learn complex nonlinear relationships, generalize across diverse network conditions and enable real-time, scalable and autonomous decision-making. Among RRM techniques, coordinated multipoint (CoMP) transmission is pivotal for mitigating inter-cell interference and enhancing cell-edge performance, thereby improving quality of experience (QoE) in dense deployments. However, optimal multi-cell selection remains a complex combinatorial challenge as it requires jointly optimizing over many possible serving-cell combinations under dynamic traffic and channel conditions. Despite their success, conventional deep reinforcement learning (DRL) methods such as proximal policy optimization (PPO) suffer from poor sample efficiency, limited generalization, and costly retraining when state and action spaces change. To address these bottlenecks, we propose a Prompt Decision Transformer (PromptDT) based multi-task learning framework capable of learning across diverse network configurations and reformulating multi-cell selection as a sequence modeling problem. By leveraging offline trajectories and task-specific prompts, PromptDT enables scalable learning across diverse network configurations, including varying base stations and user equipment counts, and scheduler policies. Experimental results demonstrate that PromptDT improves QoE by up to 49% in multi-task settings compared to baselines, with performance scaling positively alongside model capacity. Moreover, PromptDT generalizes effectively to unseen tasks, achieving robust few-shot adaptation to new network configurations without retraining or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。