arXiv:2409.15346cs.IR2024-09中稿 · publication, Accep…被引 1

用词邻域结构与相似度分析,挖掘大数据搜索中的异常行为

Big data searching using words

  • 构建词在大数据搜索中的邻域结构,为拓扑分析奠基
  • 结合杰卡德相似系数,发现搜索行为中的异常模式
  • 适合从事数据安全、搜索优化的研究者参考

大数据分析是计算机科学、企业、电子商务和国防领域最具前景的研究与开发方向之一。对众多组织而言,大数据被视为最重要的战略资产。这种爆炸式增长促使必须从数学角度发展有效的数据分析技术。在多种大数据分析方法中,拓扑数据分析(TDA)被认为是重要工具之一。然而,当前缺乏与大数据拓扑结构相关的基础概念。本文提出大数据搜索中词的邻域结构的基础概念,为未来构建大数据拓扑框架奠定基础。同时引入大数据原初(big data primal)的概念,探讨邻域结构与杰卡德相似系数结合,如何用于检测搜索行为中的异常。

原文摘要 · Abstract (English)

Big data analytics is one of the most promising areas of new research and development in computer science, enterprises, e-commerce, and defense. For many organizations, big data is considered one of their most important strategic assets. This explosive growth has made it necessary to develop effective techniques for examining and analyzing big data from mathematical perspectives. Among various methods of analyzing big data, topological data analysis (TDA) is now considered one of the useful tools. However, there is no fundamental concept related to the topological structure in big data. In this paper, we present fundamental concepts related to the neighborhood structures of words in big data search, laying the groundwork for developing topological frameworks for big data in the future. We also introduce the notion of big data primal within the context of big data search and explore how neighborhood structures, combined with the Jaccard similarity coefficient, can be utilized to detect anomalies in search behavior.

大数据分析拓扑数据异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。