提出高效采样算法,快速发现大数据中高价值关联模式。
Scalable Sampling for High Utility Patterns
- 基于新定理设计采样算法,兼顾交互性与统计可靠性。
- 在秒级内完成大规模数据的代表性模式发现。
- 适合需要快速探索高价值模式的研究者与分析师。
从数据中挖掘有意义的关联以获取有价值洞察是关键任务。然而,在定量数据库中识别代表性模式面临挑战,尤其在大规模数据下,枚举方法因搜索空间过大而效率低下。输出空间采样方法因其降低计算开销的能力成为有前景的解决方案,但现有方法在处理大型定量数据库时仍存在可扩展性问题。本文提出一种新型高实用模式采样算法及其磁盘版本,基于两个原创定理设计,适用于大规模定量数据库。该方法在保证用户中心交互性的同时,提供强大的统计保障。通过随机采样,用户可在数秒内发现相关且具代表性的高实用性模式,实现高效数据库探索。我们通过考古知识图谱子轮廓发现的典型案例展示了方法的价值。在语义与非语义定量数据库上的实验表明,本方法优于当前最优技术。
原文摘要 · Abstract (English)
Discovering valuable insights from data through meaningful associations is a crucial task. However, it becomes challenging when trying to identify representative patterns in quantitative databases, especially with large datasets, as enumeration-based strategies struggle due to the vast search space involved. To tackle this challenge, output space sampling methods have emerged as a promising solution thanks to its ability to discover valuable patterns with reduced computational overhead. However, existing sampling methods often encounter limitations when dealing with large quantitative database, resulting in scalability-related challenges. In this work, we propose a novel high utility pattern sampling algorithm and its on-disk version both designed for large quantitative databases based on two original theorems. Our approach ensures both the interactivity required for user-centered methods and strong statistical guarantees through random sampling. Thanks to our method, users can instantly discover relevant and representative utility pattern, facilitating efficient exploration of the database within seconds. To demonstrate the interest of our approach, we present a compelling use case involving archaeological knowledge graph sub-profiles discovery. Experiments on semantic and none-semantic quantitative databases show that our approach outperforms the state-of-the art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。