arXiv:2604.13271cs.LG2026-04被引 1

提升电信大模型置信度估计可靠性,避免盲目自信。

Enhancing Confidence Estimation in Telco LLMs via Twin-Pass CoT-Ensembling

  • 采用双轮独立推理+思维链集成,融合多份判断生成更准置信度。
  • 在多个电信基准上将期望校准误差降低88%,显著改善置信度可信度。
  • 适合需要高可靠输出的电信系统开发与模型评估场景。

大型语言模型(LLMs)在电信领域应用日益广泛,涵盖3GPP规范分析与O-RAN网络故障排查等复杂任务。然而,其生成的置信度分数常存在系统性偏高,缺乏可信自评能力,难以验证输出结果并安全使用。本文以代表性的Gemma-3系列模型(4B、12B、27B参数)为基础,在TeleQnA、ORANBench和srsRANBench三个基准上研究电信领域置信度校准问题。实验表明,标准单轮、口语化置信度估计无法反映真实正确性,常对错误预测赋予过高置信。为此,提出一种新型双轮思维链(CoT)集成方法,通过多次独立推理评估并聚合结果,生成校准后的置信度。该方法在多个基准上将期望校准误差(ECE)降低最高达88%,显著提升模型自我评估的可靠性。研究揭示了现有置信度估计方法的局限,并为电信领域大模型输出可信评估提供了可行路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied to complex telecommunications tasks, including 3GPP specification analysis and O-RAN network troubleshooting. However, a critical limitation remains: LLM-generated confidence scores are often biased and unreliable, frequently exhibiting systematic overconfidence. This lack of trustworthy self-assessment makes it difficult to verify model outputs and safely rely on them in practice. In this paper, we study confidence calibration in telecom-domain LLMs using the representative Gemma-3 model family (4B, 12B, and 27B parameters), evaluated on TeleQnA, ORANBench, and srsRANBench. We show that standard single-pass, verbalized confidence estimates fail to reflect true correctness, often assigning high confidence to incorrect predictions. To address this, we propose a novel Twin-Pass Chain of Thought (CoT)-Ensembling methodology for improving confidence estimation by leveraging multiple independent reasoning evaluations and aggregating their assessments into a calibrated confidence score. Our approach reduces Expected Calibration Error (ECE) by up to 88% across benchmarks, significantly improving the reliability of model self-assessment. These results highlight the limitations of current confidence estimation practices and demonstrate a practical path toward more trustworthy evaluation of LLM outputs in telecommunications.

大模型置信度估计电信AI思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。