用可执行环境评估大模型在真实网络配置中的表现
NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration

- 构建模拟多设备网络环境,支持闭环配置任务
- 480个任务实例中发现失败不仅因命令错误,还因规划不足
- 适合研究网络自动化、大模型可靠性的团队
大语言模型(LLM)代理在自动化网络配置中日益受到关注,但其可靠性与失效模式仍不明确。现有基准存在缺陷:多将配置视为静态命令生成,或依赖过于简化的设置,难以反映协议复杂性和拓扑依赖性带来的核心挑战。我们提出NetConfArena,一个面向闭环网络配置的可执行基准。该基准在模拟的多设备网络中运行代理,提供标准化动作接口,并通过隐藏的任务专属可执行测试用例评估网络行为。基准基于LLM辅助、仿真驱动的流程,将人类可读的网络材料转化为可复用的参数化任务模板。我们在96个协议聚焦的任务模板上生成480个任务实例,共获得3840条执行轨迹。结果显示,失败不仅源于命令错误,还暴露了任务规范遵循度不足及鲁棒规划与执行能力缺失。这些发现提示两个未来方向:利用验证过的执行轨迹作为监督信号改进基础模型,以及设计提升代理执行可靠性和可问责性的机制。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are poorly understood. An essential prerequisite is to assess such agents in a realistic but risk-free environment. Existing benchmarks, however, fall short: they often treat configuration as static command generation or rely on overly simplified settings. Such evaluations understate the core challenges of network configuration, where correctness requires reasoning about protocol complexity and topology dependence. We present NetConfArena, an executable benchmark for evaluating LLM agents in closed-loop network configuration. NetConfArena places agents in emulated multi-device networks, provides a standardized and compact action interface for task execution, and evaluates the resulting network behavior with hidden task-specific executable test cases. The benchmark relies on an LLM-assisted, emulation-grounded pipeline, which converts human-oriented network materials into reusable parameterized task templates. We evaluate representative LLM agents on 480 task instances instantiated from 96 protocol-focused task templates, yielding 3840 execution trajectories, and show that failures are not limited to command errors. The failures also reveal gaps in task-specification adherence and robust planning and execution. These findings suggest two future directions: using validated trajectories as supervision signals to improve foundation models, and designing harness mechanisms that make agent execution more reliable and accountable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。