arXiv:2506.00576cs.LGcs.AI2025-06被引 13

用大模型生成智能提示,提升无线网络切片的强化学习效率

ORAN-GUIDE: RAG-Driven Prompt Learning for LLM-Augmented Reinforcement Learning in O-RAN Network Slicing

  • 用O-RAN数据预训练大模型生成语义提示,增强状态表示
  • 相比基线方法,采样效率提升40%,策略收敛速度更快
  • 适合研究无线网络智能控制与大模型融合的学者

先进无线网络需应对高度动态和异构的服务需求。开放无线接入网(O-RAN)通过模块化、解耦的组件(如RAN智能控制器RIC、集中式单元CU、分布式单元DU)实现灵活性,并支持基于机器学习的智能控制。尽管深度强化学习(DRL)在动态资源分配和网络切片管理中表现强大,但常难以处理如射频特征、QoS指标和流量趋势等原始非结构化输入,导致策略泛化能力弱、决策效率低,尤其在部分可观测且持续演化的环境中。为此,我们提出ORAN-GUIDE,一种双大模型框架,通过任务相关的语义增强状态表示来提升多智能体强化学习(MARL)性能。该架构采用在O-RAN控制与配置数据上预训练的领域专用语言模型ORANSight,生成结构化、上下文感知的提示。这些提示与可学习的标记融合后,输入到冻结的GPT编码器,输出高层语义表征供DRL智能体使用。该设计采用针对无线系统技术决策优化的检索增强生成(RAG)流程。实验表明,与标准MARL及单大模型基线相比,ORAN-GUIDE显著提升了样本效率、策略收敛速度与性能泛化能力。

原文摘要 · Abstract (English)

Advanced wireless networks must support highly dynamic and heterogeneous service demands. Open Radio Access Network (O-RAN) architecture enables this flexibility by adopting modular, disaggregated components, such as the RAN Intelligent Controller (RIC), Centralized Unit (CU), and Distributed Unit (DU), that can support intelligent control via machine learning (ML). While deep reinforcement learning (DRL) is a powerful tool for managing dynamic resource allocation and slicing, it often struggles to process raw, unstructured input like RF features, QoS metrics, and traffic trends. These limitations hinder policy generalization and decision efficiency in partially observable and evolving environments. To address this, we propose \textit{ORAN-GUIDE}, a dual-LLM framework that enhances multi-agent RL (MARL) with task-relevant, semantically enriched state representations. The architecture employs a domain-specific language model, ORANSight, pretrained on O-RAN control and configuration data, to generate structured, context-aware prompts. These prompts are fused with learnable tokens and passed to a frozen GPT-based encoder that outputs high-level semantic representations for DRL agents. This design adopts a retrieval-augmented generation (RAG) style pipeline tailored for technical decision-making in wireless systems. Experimental results show that ORAN-GUIDE improves sample efficiency, policy convergence, and performance generalization over standard MARL and single-LLM baselines.

强化学习网络切片大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。