arXiv:2511.17417cs.SEcs.LG2025-11

针对电信故障报告检索难题,提出按标准分域建模的新方法,提升准确率与可解释性。

CREST: Improving Interpretability and Effectiveness of Troubleshooting at Ericsson through Criterion-Specific Trouble Report Retrieval

  • 按故障标准分域训练专用模型,集成输出增强检索信号
  • 在关键指标上显著优于单一模型,提升检索准确率与分数校准
  • 适合需要快速定位故障的运维团队和系统维护人员

通信行业快速发展要求高效故障排查以保障网络可靠性、软件可维护性和服务质量。故障报告(TRs)记录了爱立信生产系统中的问题,在及时解决软件缺陷中起关键作用。然而,TR数据的复杂性与体量,以及涵盖故障不同方面的多样标准,给检索系统带来挑战。基于爱立信前期两阶段流程(初始检索与重排序),本研究探讨不同TR观察标准对检索模型性能的影响。提出CREST(基于标准的检索集成专用报告模型),通过为不同TR领域训练专用模型并聚合输出,提升检索有效性与可解释性,支持更快故障定位与软件维护。该方法利用特定标准训练的模型捕捉多样化互补信号,显著提升检索精度、改善预测分数校准,并通过各标准相关性评分增强可解释性。基于爱立信内部TR子集的实验表明,分标准模型显著优于单一模型,验证了所有目标标准在优化检索系统中的重要性。

原文摘要 · Abstract (English)

The rapid evolution of the telecommunication industry necessitates efficient troubleshooting processes to maintain network reliability, software maintainability, and service quality. Trouble Reports (TRs), which document issues in Ericsson's production system, play a critical role in facilitating the timely resolution of software faults. However, the complexity and volume of TR data, along with the presence of diverse criteria that reflect different aspects of each fault, present challenges for retrieval systems. Building on prior work at Ericsson, which utilized a two-stage workflow, comprising Initial Retrieval (IR) and Re-Ranking (RR) stages, this study investigates different TR observation criteria and their impact on the performance of retrieval models. We propose \textbf{CREST} (\textbf{C}riteria-specific \textbf{R}etrieval via \textbf{E}nsemble of \textbf{S}pecialized \textbf{T}R models), a criterion-driven retrieval approach that leverages specialized models for different TR fields to improve both effectiveness and interpretability, thereby enabling quicker fault resolution and supporting software maintenance. CREST utilizes specialized models trained on specific TR criteria and aggregates their outputs to capture diverse and complementary signals. This approach leads to enhanced retrieval accuracy, better calibration of predicted scores, and improved interpretability by providing relevance scores for each criterion, helping users understand why specific TRs were retrieved. Using a subset of Ericsson's internal TRs, this research demonstrates that criterion-specific models significantly outperform a single model approach across key evaluation metrics. This highlights the importance of all targeted criteria used in this study for optimizing the performance of retrieval systems.

故障排查信息检索可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。