arXiv:2511.20667cs.CL2025-11被引 2

用中心点融合语义与词法特征,提升IT服务分类效率与可解释性。

A centroid based framework for text classification in itsm environments

  • 分设语义与词法中心点,推理时用倒数排名融合
  • 在123类8968条工单上达到0.731的层次化F1,训练快5.9倍
  • 适合重视可解释性与快速更新的生产级IT服务系统

在IT服务管理(ITSM)系统中,对支持工单进行层次化分类是基本需求。本文提出一种双嵌入中心点分类框架,为每个类别分别维护语义与词法中心点表示,并在推理时通过倒数排名融合(reciprocal rank fusion)结合二者。该方法在123个类别、共8,968条工单的ITSM数据集上,实现与支持向量机相当的层次化F1得分(0.731对比0.727),同时提供可解释性。相比基准方法,训练速度提升5.9倍,增量更新最快达152倍。在排除嵌入计算的前提下,批处理规模(100-1000样本)下整体提速8.6至8.8倍。这些性能使该方法适用于强调可解释性与运行效率的生产级ITSM环境。

原文摘要 · Abstract (English)

Text classification with hierarchical taxonomies is a fundamental requirement in IT Service Management (ITSM) systems, where support tickets must be categorized into tree-structured taxonomies. We present a dual-embedding centroid-based classification framework that maintains separate semantic and lexical centroid representations per category, combining them through reciprocal rank fusion at inference time. The framework achieves performance competitive with Support Vector Machines (hierarchical F1: 0.731 vs 0.727) while providing interpretability through centroid representations. Evaluated on 8,968 ITSM tickets across 123 categories, this method achieves 5.9 times faster training and up to 152 times faster incremental updates. With 8.6-8.8 times speedup across batch sizes (100-1000 samples) when excluding embedding computation. These results make the method suitable for production ITSM environments prioritizing interpretability and operational efficiency.

文本分类IT服务管理可解释性中心点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。