arXiv:2608.26126cs.CLcs.IT2026-08

开源电信领域推理模型,能精准处理协议、故障等多类专业问题。

TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

论文配图:TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack
图 1 · 摘自论文原文
  • 构建四轴数据集,融合协议、知识、建模与故障推理
  • 在7个公开电信基准上超越其他开源模型,接近闭源顶尖水平
  • 适合电信工程师、AI研究者及自动化运维系统开发者

电信是大语言模型推理的高价值领域,因工程流程需结合规范文档、运行数据、厂商故障信息与精确射频/网络计算。然而当前电信LLM应用受限于双侧能力缺口:通用推理模型缺乏领域专精,而专用模型又欠缺结构化多步推理能力。为此,我们发布TelecomGPT-R1-9B,一个统一的开源电信推理模型,在GSMA开放电信排行榜中表现领先。我们构建了包含67,427条样本的监督微调(SFT)语料库,围绕协议、知识、建模和故障四个互补推理维度组织,数据来自匹配领域的公开网络资源,并通过轴向链式思维生成与前缀续写自验证增强。基于Qwen3.5-9B,采用两阶段后训练方案:第一阶段使用多教师低秩适配(LoRA)SFT注入电信知识并诱导轴向推理格式;第二阶段采用组相对策略优化(GRPO),通过解耦裁剪与动态采样策略优化(DAPO)稳定训练,利用四个轴对齐的二元验证器奖励优化策略。在七个公开电信基准测试中,TelecomGPT-R1-9B在开源模型中排名第一,七轴平均表现接近最先进的闭源前沿推理模型。

原文摘要 · Abstract (English)

Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervised fine-tuning (SFT) corpus organized around four complementary reasoning axes: protocol, knowledge, modeling, and fault. The corpus is built from axis-matched public web sources and enhanced through axis-specific chain-of-thought (CoT) generation and prefix-continuation self-validation. Starting from Qwen3.5-9B, we further develop a two-stage post-training recipe. First, multi-teacher low-rank adaptation (LoRA)-based SFT injects telecom knowledge and induces axis-specific reasoning formats. Second, group relative policy optimization (GRPO), stabilized by decoupled clip and dynamic sampling policy optimization (DAPO), optimizes the policy using four axis-aligned binary verifier rewards. Across seven public telecom benchmarks, TelecomGPT-R1-9B ranks first among open-source telecom LLMs and achieves a seven-axis mean comparable to state-of-the-art closed-source frontier reasoners.

电信推理开源模型多步推理知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。