arXiv:2511.03107cs.LGcs.IT2025-11

用改进的TF-IDF和快速降维算法,让传统模型更省电更快。

An Efficient Classification Model for Cyber Text

  • 提出CTF-IDF改进文本特征提取,提升信息密度。
  • 结合IRLBA降维,使模型训练时间大幅缩短。
  • 适合追求低功耗、快响应的文本分类场景。

近年来深度学习迅猛发展,但其对计算资源和电力的高需求带来了严重的碳足迹问题。文本分析领域也深受这一趋势影响。本文提出了一种改进的TF-IDF算法——克莱门特词频-逆文档频率(CTF-IDF),用于数据预处理,并结合一种快速的IRLBA算法进行降维。将这两项技术引入传统文本分析流程,相比深度学习方法,在显著降低计算开销与碳排放的同时,仅以轻微精度损失为代价,实现了更高效、更快的分类应用。实验结果表明,所提方法在时间复杂度上实现数量级降低,且模型准确率进一步提升。

原文摘要 · Abstract (English)

The uprising of deep learning methodology and practice in recent years has brought about a severe consequence of increasing carbon footprint due to the insatiable demand for computational resources and power. The field of text analytics also experienced a massive transformation in this trend of monopolizing methodology. In this paper, the original TF-IDF algorithm has been modified, and Clement Term Frequency-Inverse Document Frequency (CTF-IDF) has been proposed for data preprocessing. This paper primarily discusses the effectiveness of classical machine learning techniques in text analytics with CTF-IDF and a faster IRLBA algorithm for dimensionality reduction. The introduction of both of these techniques in the conventional text analytics pipeline ensures a more efficient, faster, and less computationally intensive application when compared with deep learning methodology regarding carbon footprint, with minor compromise in accuracy. The experimental results also exhibit a manifold of reduction in time complexity and improvement of model accuracy for the classical machine learning methods discussed further in this paper.

文本分类降维低碳

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。