arXiv:2605.18025cs.AI2026-05KDD被引 2

构建电信领域评测基准,揭示大模型在工业应用中的执行短板

TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications?

论文配图:TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications?
图 1 · 摘自论文原文
  • 基于知识图谱融合通信原理、3GPP协议与产商专有知识,评估多维度理解能力
  • 覆盖6类真实网络运维流程,模型在意图识别准确率达90%,但解决方案生成仅30%
  • 揭示大模型作为诊断员可行,但无法胜任现场工程师的现实困境,适合垂域优化研究者

尽管大语言模型已在多个垂直领域取得显著进展,但在电信领域的部署仍处于探索阶段,主要受限于缺乏标准化评估框架。现有电信基准多聚焦静态基础认知与孤立技能,忽视设备专属文档和端到端工业工作流。为此,我们提出TeleCom-Bench,一个包含12个评估集、共22,678个样本的综合性基准,涵盖两个协同层级:(1)多维知识理解,通过知识图谱驱动融合通信基础、3GPP协议、5G网络架构及有线、核心、无线网络的专有产品知识;(2)端到端知识应用,形式化六项核心任务——意图识别、实体抽取、事件验证、工具调用、根因分析、解决方案生成——基于真实网络运维代理工作流,覆盖网络优化与故障维护场景。对八种先进大模型的评估显示普遍存在“执行墙”:模型在语言接口任务(如意图识别、实体抽取)上可达90%准确率,但在流程执行任务(如解决方案生成)中骤降至约30%。这一能力差距表明当前模型仅能胜任诊断角色,难以担当现场工程师。TeleCom-Bench提供标准化诊断工具,精准定位此缺陷,为打造生产级电信智能体提供可操作指引。数据集与评估代码已开源至https://github.com/ZTE-AICloud/TeleCom-Bench。

原文摘要 · Abstract (English)

While Large Language Models have achieved remarkable integration in various vertical scenarios, their deployment in the telecommunications domain remains exploratory due to the lack of a standardized evaluation framework. Current telecom benchmarks primarily focus on static, foundational knowledge and isolated atomic skills, neglecting the equipment-specific documentation and end-to-end industrial workflows essential for real-world production systems. To bridge this gap, we present TeleCom-Bench, a comprehensive benchmark comprising 12 evaluation sets with 22,678 curated samples, which evaluates LLMs across a synergistic hierarchy: (1) Multi-dimensional Knowledge Comprehension, which integrates telecommunication fundamentals, 3GPP protocols, and 5G network architecture with proprietary product knowledge across wired, core, and wireless networks via knowledge graph-driven synthesis; and (2)End-to-End Knowledge Application, which formalizes six core tasks on authentic trajectories from live network agent workflows, including intent recognition, entity extraction, event verification, tool invocation, root cause analysis, and solution generation-across network optimization and fault maintenance scenarios. Evaluations of eight state-of-the-art LLMs reveal a universal Execution Wall: while models achieve 90% accuracy in linguistic interface tasks such as intent recognition and entity extraction, performance collapses to approximately 30% in procedural execution tasks like solution generation. This capability gap demonstrates that current LLMs function competently as diagnosticians but fail as field engineers. TeleCom-Bench provides standardized diagnostics to precisely pinpoint this deficit, offering actionable guidance for domain-specific alignment toward production-ready telecom agents. The dataset and evaluation code have been released at https://github.com/ZTE-AICloud/TeleCom-Bench.

大模型评测电信智能化知识图谱工业落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。