构建可复现的多厂商配置翻译语义基准,揭示LLM生成配置的隐藏缺陷。
A Reproducible Semantic Benchmark for Multivendor DSM-to-CLI Translation

- 设计覆盖5大云LLM、3厂商、5用例的可复现语义评测框架
- 发现语义质量与运行可靠性正交,华为设备暴露被聚合指标掩盖的失败模式
- 重复实验波动性可预测投票不稳定性,适合评估多厂商网络自动化系统
将高层网络意图转化为正确多厂商配置仍是网络自动化的核心挑战,因语法正确的输出仍可能违背预期运行状态。尽管大语言模型(LLMs)取得进展,领域仍缺乏可复现的语义基准用于严格跨厂商评估。本文提出一个可复现的DSM-to-CLI语义基准,涵盖五种云LLM、三家厂商、五个典型用例,每组实验重复十次,采用固定评审员和明确的失败分类体系。结果表明:语义质量与操作可靠性相互独立;厂商影响显著大于用例影响;重复实验的分散性强烈预示投票不稳定性,华为VRP暴露了聚合指标无法捕捉的故障模式。研究证明,多厂商、重复执行的语义基准对科学评估基于LLM的网络配置系统至关重要。
原文摘要 · Abstract (English)
Translating high-level network intents into correct multivendor configurations remains a central challenge in network automation, as syntactically valid outputs may still violate the intended operational state. Despite recent advances in Large Language Models (LLMs), the field still lacks reproducible semantic benchmarks for rigorous cross-vendor evaluation. This paper presents a reproducible DSM-to-CLI semantic benchmark covering five cloud LLMs, three vendors, five representative use cases, and ten repeated runs per experimental cell under fixed judges and an explicit failure taxonomy. Our results show that semantic quality and operational reliability are orthogonal, vendor effects dominate use-case effects, and repeated-run dispersion strongly predicts vote instability, with Huawei VRP exposing failure modes hidden by aggregate metrics. These findings demonstrate that multivendor, repeated-execution semantic benchmarks are essential for scientifically rigorous comparison of LLM-based network configuration systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。