arXiv:2608.21021cs.CLcs.AI2026-08中稿 · presentation in IE…

轻量级大模型可高效完成5G故障分析,但对规范记忆仍不足。

Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

  • 用自由文本评估轻量级大模型在5G领域的诊断能力
  • 三模型故障诊断准确率均超90%,但规范召回率低于60%
  • 多模型评分一致性强,适合工业级部署

5G及未来6G网络的真实故障分析需依赖领域知识解析自由文本诊断内容,包括根本原因解释与处理建议。尽管大语言模型(LLM)在自动化该任务中展现潜力,但轻量级边缘部署模型是否具备深度文本诊断能力仍不明确。现有基准多采用封闭式选择题,而本文首次在自由文本生成范式下评估3个轻量级LLM:Claude-Haiku-4.5、GPT-5.4-Mini与Gemini-3.1-Flash-Lite,覆盖TeleQNA ORAN FT、5G-Faults FT和TeleInter FT三个基准。通过三位前沿专家级裁判评分,验证了多裁判评分的一致性。所有模型故障诊断准确率均达90%以上,但对3GPP与O-RAN规范的零样本召回率均低于60%。跨轮次平均裁判间一致性系数(ICC)不低于0.90,证明基于多裁判的LLM评分框架可生成稳定可靠的开放文本评价结果。实际部署中,Gemini-3.1-Flash-Lite在精度与推理成本/延迟间表现最佳,是最适于生产环境的候选模型。

原文摘要 · Abstract (English)

Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating this, yet whether lightweight, edge-deployable models are capable of performing in-depth free-text diagnostics remains an open question. While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format. Transitioning to this paradigm requires evaluating lightweight, edge-deployable AI models on open-ended diagnostic reasoning, alongside a dependable framework to validate these text outputs at scale. To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter FT. Three independent frontier judges score outputs, and pairwise inter-judge agreement is measured as an empirical test of the LLM-as-Judge methodology. All three models reach at least 90% accuracy on fault diagnosis, while zero-shot recall of 3GPP and O-RAN specifications remains the critical gap, with all models scoring below 60%. Mean inter-judge agreement is at least 0.90 across all runs, indicating that multi-judge LLM scoring produces consistent, reproducible grades for open-ended telecom responses. Operationally, Gemini-3.1-Flash-Lite offers the best efficiency trade-off, combining competitive accuracy with the lowest inference cost and latency, making it the most suitable candidate for production telecom deployments.

5G故障分析大模型评估轻量级模型自由文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。