arXiv:2605.15110cs.LGcs.CL2026-05被引 3

提出统计特征提升字符串相似度计算,不依赖语言信息。

Proposal and study of statistical features for string similarity computation and classification

  • 用共现矩阵和游程矩阵计算字符串相似性
  • 在合成与真实数据上均优于现有方法
  • 适合跨语言、无语法结构的文本分析

针对通用字符串(单词、短语、代码、文本)的相似度计算,提出了源自视觉计算领域的共现矩阵(COM)和游程矩阵(RLM)特征。这些特征完全基于统计,不依赖语言信息,适用于任意语言和语法结构。对比了最长公共子序列、最大连续最长公共子序列、互信息和编辑距离等常用统计度量。在首组合成实验中,COM和RLM特征表现最优;在4个案例中有3个,其显著性高于第二优的度量组(P值<0.001)。在真实文本抄袭检测数据集上,RLM特征取得最佳效果。

原文摘要 · Abstract (English)

Adaptations of features commonly applied in the field of visual computing, co-occurrence matrix (COM) and run-length matrix (RLM), are proposed for the similarity computation of strings in general (words, phrases, codes and texts). The proposed features are not sensitive to language related information. These are purely statistical and can be used in any context with any language or grammatical structure. Other statistical measures that are commonly employed in the field such as longest common subsequence, maximal consecutive longest common subsequence, mutual information and edit distances are evaluated and compared. In the first synthetic set of experiments, the COM and RLM features outperform the remaining state-of-the-art statistical features. In 3 out of 4 cases, the RLM and COM features were statistically more significant than the second best group based on distances (P-value < 0.001). When it comes to a real text plagiarism dataset, the RLM features obtained the best results.

字符串相似统计特征文本检测无语言依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。