打造专用于电信领域的开源大模型基础,提升智能网络任务表现
OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks

- 构建统一电信数据资源,涵盖检索、重排序与指令微调
- 三大模型性能显著提升:检索93.5% NDCG@10,重排序0.952 MRR@10
- 适合电信AI研究者与从业者,助力构建专业大模型
前沿AI模型虽发展迅速,但在电信领域任务上仍表现不佳。本文提出开放电信人工智能资源Open Telco(OTel),包含检索、重排序、指令微调及安全/回避数据集,并提供30个全参数后训练基线模型,覆盖嵌入、重排序与语言模型三类。基于已有开放电信数据集与基准,OTel整合了可复现的电信数据源、预留评估划分、训练好的嵌入模型、重排序器、上下文感知大模型及安全/回避数据。后训练使三类模型性能全面提升:嵌入检索达到93.5% NDCG@10,重排序达0.952 MRR@10,语言模型正确率88.2%。截至2026年5月3日,项目模型下载量超1600万次,获全球157+媒体报道。我们开源该资源,欢迎社区扩展数据、优化模型并构建更强的上下文感知电信大模型。
原文摘要 · Abstract (English)
Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open telecom AI resource with derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, plus 30 full-parameter post-trained baselines across embedding, reranking, and language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times, and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.5% NDCG@10, reranking reaches 0.952 MRR@10, and language-model correctness reaches 88.2%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。