arXiv:2512.00105cs.DBcs.AI2025-12被引 1

提出高效采样数值数据库中区间模式的新方法,解决长尾问题。

Efficiently Sampling Interval Patterns from Numerical Databases

  • 基于频率比例多步采样,精准统计覆盖对象的区间模式数。
  • Fips按频率采样,HFips结合频率与超体积双重加权,提升代表性。
  • 适合数据挖掘、异常检测等需高效发现关键模式的场景。

模式采样已成为大型数据库中信息发现的有力手段,使分析者能聚焦于可管理的模式子集。本文首次提出针对数值数据库中区间模式的采样方法Fips,该方法按模式频率进行比例采样。Fips采用多步采样流程,解决了数值数据中的核心挑战:准确计算每个对象被多少区间模式覆盖。进一步提出HFips,使采样比例同时依赖于模式频率和超体积(hyper-volume)。这些方法有效应对了模式采样中普遍存在的长尾现象。我们形式化证明了Fips按频率采样,HFips按频率与超体积乘积采样。在多个数据集上的实验表明,所获模式质量高,且对长尾现象具有鲁棒性。

原文摘要 · Abstract (English)

Pattern sampling has emerged as a promising approach for information discovery in large databases, allowing analysts to focus on a manageable subset of patterns. In this approach, patterns are randomly drawn based on an interestingness measure, such as frequency or hyper-volume. This paper presents the first sampling approach designed to handle interval patterns in numerical databases. This approach, named Fips, samples interval patterns proportionally to their frequency. It uses a multi-step sampling procedure and addresses a key challenge in numerical data: accurately determining the number of interval patterns that cover each object. We extend this work with HFips, which samples interval patterns proportionally to both their frequency and hyper-volume. These methods efficiently tackle the well-known long-tail phenomenon in pattern sampling. We formally prove that Fips and HFips sample interval patterns in proportion to their frequency and the product of hyper-volume and frequency, respectively. Through experiments on several databases, we demonstrate the quality of the obtained patterns and their robustness against the long-tail phenomenon.

模式采样数值数据长尾现象区间模式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。