arXiv:2603.10564cs.AIcs.NI2026-03

让AI自主优化网络切片,无需人工设定奖励信号。

Adaptive RAN Slicing Control via Reward-Free Self-Finetuning Agents

  • 用自我反思生成反馈,把经验直接融入模型参数。
  • 在动态网络中同时优化速率、质量与稳定性,性能超越传统强化学习。
  • 适合研究智能网络控制或大模型持续学习的读者。

将生成式AI引入AI原生网络系统,为实现自主自适应控制提供了新路径。然而,现有模型在连续控制任务中受限于有限上下文窗口、缺乏显式奖励信号及长上下文退化等问题。本文提出一种新型自微调框架,使智能体通过直接环境交互实现持续学习,无需人工设计奖励。框架采用双视角反思机制,从交互历史中自动生成语言反馈,构建偏好数据集;随后通过基于偏好的微调,将长时序经验固化到模型参数中。我们在动态无线接入网(RAN)切片任务上评估该方法,这是一个需在波动网络条件下权衡频谱效率、服务质量与重配置稳定性的复杂多目标控制问题。实验表明,该框架在样本效率、稳定性及多指标优化方面均优于标准强化学习基线和现有大语言模型代理。结果验证了自进化生成式智能体在连续控制中的潜力,为未来AI原生网络基础设施奠定基础。

原文摘要 · Abstract (English)

The integration of Generative AI models into AI-native network systems offers a transformative path toward achieving autonomous and adaptive control. However, the application of such models to continuous control tasks is impeded by intrinsic architectural limitations, including finite context windows, the lack of explicit reward signals, and the degradation of the long context. This paper posits that the key to unlocking robust continuous control is enabling agents to internalize experience by distilling it into their parameters, rather than relying on prompt-based memory. To this end, we propose a novel self-finetuning framework that enables agentic systems to learn continuously through direct interaction with the environment, bypassing the need for handcrafted rewards. Our framework implements a bi-perspective reflection mechanism that generates autonomous linguistic feedback to construct preference datasets from interaction history. A subsequent preference-based fine-tuning process distills long-horizon experiences into the model's parameters. We evaluate our approach on a dynamic Radio Access Network (RAN) slicing task, a challenging multi-objective control problem that requires the resolution of acute trade-offs between spectrum efficiency, service quality, and reconfiguration stability under volatile network conditions. Experimental results show that our framework outperforms standard Reinforcement Learning (RL) baselines and existing Large Language Model (LLM)-based agents in sample efficiency, stability, and multi-metric optimization. These findings demonstrate the potential of self-improving generative agents for continuous control tasks, paving the way for future AI-native network infrastructure.

网络优化生成式AI持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。