构建动态基准,评估网络智能体多轮操作的稳定性
NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration
- 用有限状态机建模网络配置流程,确保行为可复现
- 四款主流大模型智能体在复杂配置中出现探索崩溃与逻辑断裂
- 适合研究自主网络系统可靠性的研究人员参考
随着智能体化网络管理日益流行,亟需超越静态单次测试的评估框架。为此,我们提出 NetAgentBench,一个基于有限状态机(FSM)形式化建模的动态基准,确保评估过程的确定性、正确性与执行边界。该框架为衡量复杂、多轮的网络操作行为提供了严谨基础。对四个前沿大语言模型智能体在多样化网络配置任务中的实证评估显示:尽管它们能完成基础任务,但在专家级配置中严重出现探索失控与连贯性崩溃。结果表明,系统性评估多轮行为稳定性,是实现可信全自治网络不可或缺的一步。
原文摘要 · Abstract (English)
As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address this, we introduce NetAgentBench, a dynamic benchmark that evaluates agent interactions through a Finite State Machine (FSM) formalization guaranteeing determinism, correctness, and bounded execution. This provides the networking landscape with a rigorous foundation to measure complex, multi-turn operational behaviors. Our empirical evaluation of four state-of-the-art LLM agents through diverse network configuration tasks reveals stark deficiencies: while agents can solve basic tasks, they suffer severe exploration meltdowns and coherence collapse during expert-level configurations. Ultimately, NetAgentBench demonstrates that systematically evaluating multi-turn behavioral stability is an indispensable step toward realizing trustworthy, fully autonomous networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。