arXiv:2603.18413stat.MLcs.LG2026-03被引 1

提出一套检验聚类流水线结果可靠性的统计方法。

Statistical Testing Framework for Clustering Pipelines by Selective Inference

  • 基于选择性推断构建聚类流水线的统计检验框架。
  • 在合成与真实数据上验证了方法能严格控制第一类错误率。
  • 适合需要可信聚类结果的数据科学家和领域研究者。

数据分析流水线是将原始数据转化为有意义洞察的结构化步骤序列,通常包含多个依赖数据的分析过程。在实际应用中,分析结果往往经过一系列数据驱动的处理步骤后才得出。本文聚焦于量化此类流水线输出结果的统计可靠性,以聚类流水线为案例,通过异常值检测、特征选择和聚类等步骤从复杂异构数据中识别聚类结构。提出一种基于选择性推断的新统计检验框架,可系统构建由预定义组件构成的聚类流水线的有效统计检验。理论证明该检验在任意名义显著性水平下均能控制第一类错误率,并在合成与真实数据集上验证其有效性和可靠性。

原文摘要 · Abstract (English)

A data analysis pipeline is a structured sequence of steps that transforms raw data into meaningful insights by integrating multiple analysis algorithms. In many practical applications, analytical findings are obtained only after data pass through several data-dependent procedures within such pipelines. In this study, we address the problem of quantifying the statistical reliability of results produced by data analysis pipelines. As a proof of concept, we focus on clustering pipelines that identify cluster structures from complex and heterogeneous data through procedures such as outlier detection, feature selection, and clustering. We propose a novel statistical testing framework to assess the significance of clustering results obtained through these pipelines. Our framework, based on selective inference, enables the systematic construction of valid statistical tests for clustering pipelines composed of predefined components. We prove that the proposed test controls the type I error rate at any nominal level and demonstrate its validity and effectiveness through experiments on synthetic and real datasets.

聚类分析选择性推断统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。