arXiv:2604.06209cs.CL2026-04被引 4

构建多语言电信AI代理评测基准,评估其理解与执行能力。

TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents

论文配图:TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents
图 1 · 摘自论文原文
  • 设计针对电信场景的多语言评测框架,聚焦意图识别与流程对齐。
  • 发现主流模型在变体场景下难以稳定遵循排查步骤,尤其在双语环境更差。
  • 适合电信运维、AI代理开发与测试人员参考,推动多语言智能运维落地。

大语言模型(LLM)代理融入电信网络带来新挑战,涉及意图识别、工具执行与解决方案生成,同时需考虑不同运营约束。本文提出面向多语言电信领域的评测框架 TelcoAgent-Bench 与 TelcoAgent-Metrics,用于评估多语言电信 LLM 代理的语义理解、流程对齐性及在重复场景变化下的稳定性。该框架包含结构化指标,涵盖意图识别、有序工具执行、结果正确性及场景变异下的稳定性,旨在量化 LLM 代理在电信环境中的可靠性与运行一致性。框架支持英文与阿拉伯语,以满足实际运维中的多语言部署需求。实验表明,尽管近期指令微调模型能合理理解电信问题,但通常难以一致遵循所需排查步骤,且在不同场景变体下行为不稳定;该性能差距在无约束和双语环境下尤为显著。

原文摘要 · Abstract (English)

The integration of large language model (LLM) agents into telecom networks introduces new challenges, related to intent recognition, tool execution, and resolution generation, while taking into consideration different operational constraints. In this paper, we introduce TelcoAgent-Bench and TelcoAgent-Metrics, a Telecom-specific benchmarking framework for evaluating multilingual telecom LLM agents. The proposed framework assesses the semantic understanding as well as process-level alignment with structured troubleshooting flows and stability across repeated scenario variations. Our contribution includes a structured suite of metrics that assess intent recognition, ordered tool execution, resolution correctness, and stability across scenario variations, with the aim of quantifying the reliability and operational consistency of LLM agents in telecom environments. The framework is designed to operate in both English and Arabic, to address the need for multilingual agent deployment in operational network environments. Our experimental results show that although recent instruct-tuned models can understand telecom problems in a reasonable way, they usually struggle to consistently follow the required troubleshooting steps and to maintain stable behavior when exposed to different variations of the same scenario. This performance gap becomes more pronounced in unconstrained and bilingual settings.

电信AI多语言评测基准LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。