构建6G网络管理闭环智能体训练框架,支持工具操作与动态学习。
6GAgentGym: Tool Use, Data Synthesis, and Agentic Learning for Network Management
- 设计42类工具的交互环境,区分观测与配置操作
- 通过仿真数据训练实验模型,实现自动生成训练轨迹
- 开源8B模型在长任务上性能逼近GPT-5,适合智能运维研究
自主6G网络管理需要能执行工具、观察状态变化并据此调整决策的智能体。现有基于静态问题或脚本回放的基准无法支持闭环交互,使智能体仅能被动评估而无法从环境反馈中学习。本文提出6GAgentGym,提供闭环交互能力:包含42种类型工具的交互环境,其效果分类区分只读观测与状态变更配置;基于NS-3仿真数据训练的实验模型可校准行为预测。6G-Forge通过迭代自指导生成并结合实验模型验证,从NS-3种子中构建闭环训练轨迹。在生成语料上进行监督微调后,再通过在线闭环强化学习,使8B开源模型在配套的6GAgentBench上达到与GPT-5相当的整体成功率,且在长时序任务上表现更优。这些组件共同为自主闭环网络管理提供了可行路径。
原文摘要 · Abstract (English)
Autonomous 6G network management requires agents that can execute tools, observe the resulting state changes, and adapt their decisions accordingly. Existing benchmarks based on static questions or scripted episode replay, however, do not support such closed-loop interaction, limiting agents to passive evaluation without the ability to learn from environmental feedback. This paper presents 6GAgentGym to provide closed-loop capability. The framework provides an interactive environment with 42 typed tools whose effect classification distinguishes read-only observation from state-mutating configuration, backed by a learned Experiment Model calibrated on NS-3 simulation data. 6G-Forge bootstraps closed-loop training trajectories from NS-3 seeds via iterative Self-Instruct generation with execution verification against the Experiment Model. Supervised fine-tuning on the resulting corpus followed by reinforcement learning with online closed-loop interaction enables an 8B open-source model to achieve comparable overall success rate to GPT-5 on the accompanying 6GAgentBench, with stronger performance on long-horizon tasks. Together, these components provide a viable path toward autonomous, closed-loop network management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。