arXiv:2506.11017cs.CLcs.AI2025-06被引 1

首个电信运维调度评估基准,测试大模型在真实场景中的表现

TeleEval-OS: Performance evaluations of large language models for operations scheduling

  • 构建15个数据集覆盖4类运维任务,分四级评估模型能力
  • 开源模型在部分任务上超越闭源模型,展现应用潜力
  • 适合关注电信AI落地与模型评测的研究者和工程师

大语言模型(LLM)的快速发展推动了人工智能在多个专业领域的应用。电信运维调度(OS)是电信行业核心环节,涉及网络、服务、风险与人力的协同管理,旨在优化生产调度并实现统一服务管控。然而,运维任务本身复杂且领域专用,加之缺乏全面的评估基准,限制了大模型在此领域的深入探索。为此,我们提出了首个电信运维调度评估基准——TeleEval-OS。该基准包含13个子任务的15个数据集,全面模拟智能工单创建、处理、关闭及评估四个关键阶段。为系统评估不同复杂度任务下的模型性能,我们将能力分为四层:基础NLP、知识问答、报告生成、报告分析。我们在TeleEval-OS上采用零样本与少样本方法,对10个开源模型(如DeepSeek-V3)和4个闭源模型(如GPT-4o)进行了多场景评估。实验表明,开源模型在特定任务中可超越闭源模型,凸显其在电信运维调度中的巨大潜力与价值。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has significantly propelled progress in artificial intelligence, demonstrating substantial application potential across multiple specialized domains. Telecommunications operation scheduling (OS) is a critical aspect of the telecommunications industry, involving the coordinated management of networks, services, risks, and human resources to optimize production scheduling and ensure unified service control. However, the inherent complexity and domain-specific nature of OS tasks, coupled with the absence of comprehensive evaluation benchmarks, have hindered thorough exploration of LLMs' application potential in this critical field. To address this research gap, we propose the first Telecommunications Operation Scheduling Evaluation Benchmark (TeleEval-OS). Specifically, this benchmark comprises 15 datasets across 13 subtasks, comprehensively simulating four key operational stages: intelligent ticket creation, intelligent ticket handling, intelligent ticket closure, and intelligent evaluation. To systematically assess the performance of LLMs on tasks of varying complexity, we categorize their capabilities in telecommunications operation scheduling into four hierarchical levels, arranged in ascending order of difficulty: basic NLP, knowledge Q&A, report generation, and report analysis. On TeleEval-OS, we leverage zero-shot and few-shot evaluation methods to comprehensively assess 10 open-source LLMs (e.g., DeepSeek-V3) and 4 closed-source LLMs (e.g., GPT-4o) across diverse scenarios. Experimental results demonstrate that open-source LLMs can outperform closed-source LLMs in specific scenarios, highlighting their significant potential and value in the field of telecommunications operation scheduling.

大模型评测电信运维智能调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。