arXiv:2605.09929cs.LGcs.SE2026-05被引 2

测试大模型在通信任务中修复错误推理的能力,发现最强模型成功率仅29.1%

TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications

论文配图:TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications
图 1 · 摘自论文原文
  • 用中断的错误推理链测试模型续写纠错能力
  • 最强模型平均修复率仅29.1%,规模提升不保证韧性增强
  • 适合关注模型鲁棒性与成本效益的通信系统开发者

将大语言模型部署于通信领域需超越任务准确率。实际流程中,模型常需承接部分完成的错误推理并继续修正。我们提出TeleResilienceBench基准,量化此能力——即推理韧性,涵盖七类电信子领域,源自GSMA Open-Telco LLM套件。数据通过收集弱生成模型的失败案例,在推理中途截断,要求目标模型续写并纠正。引入正确翻转率(CFR)作为恢复成功度量,评估了来自Qwen3.5、Gemma4和Nemotron-3系列的八款模型。结果表明,即使最强模型宏观平均CFR也仅为29.1%,且同系列内规模提升无法可靠改善韧性。Nemotron-3-nano 4b优于所有Qwen3.5变体(含27b),在辅助的TeleMath数值评估中达23.4% CR%,展现出最佳韧性-成本比。难度分层分析显示,现有电信基准难度标签反映事实特异性而非推理深度,说明当前评估更侧重知识覆盖而非推理能力。

原文摘要 · Abstract (English)

Deploying large language models in telecommunications requires more than task accuracy. In realistic workflows, a model may inherit partially completed reasoning from a prior step, an upstream agent, or its own earlier generation, and must continue that reasoning even when it is already going wrong. We introduce TeleResilienceBench, a benchmark that quantifies this capability, which we term reasoning resilience, across seven telecom sub-domains drawn from the GSMA Open-Telco LLM suite. Instances are constructed by collecting failures from a weak generator model, truncating the flawed reasoning trace at its midpoint, and asking a target model to continue and correct it. We propose the Correct Flip Rate (CFR) as a direct measure of successful recovery and evaluate eight models spanning the Qwen3.5, Gemma4, and Nemotron-3 families. Our results show that even the strongest model achieves a macro-average CFR of only 29.1%, and scale does not reliably improve resilience within families. Nemotron-3-nano 4b outperforms all Qwen3.5 variants including the 27b model and leads the auxiliary TeleMath numerical evaluation at 23.4% CR%, offering the best resilience-to-cost ratio in the set. A difficulty-stratified analysis further reveals that existing telecom benchmark difficulty labels reflect factual specificity rather than reasoning depth, suggesting that current evaluations measure knowledge coverage more than reasoning ability.

大模型推理韧性通信基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。