提出少样本重采样方法FewRS,显著降低统计验证计算开销。
Few-Shot Resampling for Scalable Statistically-Sound Data Mining
- 基于新推导的统计量上界,仅需极少重采样数据集
- 相比现有方法提速最高达100倍,仍保持高检验效能
- 适合大规模数据挖掘结果的可靠统计验证
知识发现的关键步骤是评估数据挖掘结果。在模式挖掘、图分析等应用中,需评估结果的统计显著性,以避免仅由噪声或数据随机波动引起的虚假发现。尽管特定场景有专用方法,但重采样方法因无需解析解而被广泛使用。然而,现有方法需生成并分析数千个重采样数据集,对大规模数据或复杂分析不切实际。本文提出FewRS,一种简单高效的重采样方法,可在严格控制假阳性概率的前提下评估数据挖掘结果的统计显著性。该方法基于对数据挖掘结果质量测试统计量上偏差的新边界推导,证明其仅需极少量重采样数据集即可实现高可扩展性与广泛应用。我们在常见任务如模式挖掘和网络分析上测试,结果表明,相比当前最优方法,运行时间最多降低两个数量级,同时保持高统计功效,实现了大规模真实世界数据集上数据挖掘结果的可靠统计验证。
原文摘要 · Abstract (English)
A key step in knowledge discovery is the evaluation of data mining results. In several applications, including pattern mining, graph analysis, and others, this step includes the evaluation of the statistical significance of the results, to avoid spurious discoveries due only to noise or random fluctuations in the data. While specialized procedures have been developed for some specific applications, resampling-based approaches are widely used, in particular for complex analyses where analytical results cannot be derived. However, current resampling-based approaches require the generation and analysis of thousands of resampled datasets, and are therefore impractical for large datasets or computationally intensive analyses. In this paper, we introduce FewRS, a simple and effective resampling-based approach to assess the statistical significance of data mining results with rigorous guarantees on the probability of false discoveries. Our approach can be used in every situation where resampling-based approaches are applied. FewRS builds on our derivation of a novel bound to the supremum deviation of test statistics representing the quality of data mining results. We prove that FewRS needs to generate and analyze an extremely small number of resampled datasets, leading to a highly scalable approach with wide applicability. We test our approach on common tasks such as pattern mining and network analysis. In all cases, our approach results in a reduction of up to two orders of magnitude in running time compared to the state of the art, while preserving high statistical power, enabling the statistical validation of data mining results on large-scale real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。