用专家评分+大模型打分,高效评估网络自动化中大模型的表现。
Human Grounded Evaluation of Large Language Models for Optical Network Automation

- 构建分级评估流程,结合专家与大模型评分
- 12B参数模型在解释质量与效率上表现最佳
- 适合需要可靠输出的网络运维自动化场景
大型语言模型(LLMs)在网络自动化中应用日益广泛,但其输出质量与推理成本在不同模型间差异显著。本文提出HuGLEN,一种分步评估流程,利用大模型作为评判者并结合少量专家评分,实现候选LLM的可扩展、可复现比较,并通过质量效率得分(QES)进行排序。我们以可解释人工智能(XAI)模型在光传输质量(QoT)估计任务中的输出为对象,验证了其转化为运维人员友好的解释能力。结果表明,一个中等规模的模型(120亿参数)在QES上表现最优,体现了解释质量与效率的最佳平衡。整体而言,HuGLEN在减少人工标注负担的同时,支持面向运维人员的自动化任务中一致的模型选择。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into operator-friendly explanations. Our results show that a medium-sized LLM (12B parameters) achieves the highest QES, indicating the best trade-off between explanation quality and efficiency. Overall, HuGLEN reduces the human-labeling burden while supporting consistent model selection for operator-facing automation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。