arXiv:2412.15894cs.LGstat.ML2024-12被引 1

无需参数设定,自动拆分多峰数据并建模。

Statistical Modeling of Univariate Multimodal Data

  • 通过密度谷点递归分割单变量数据为单峰子集。
  • 在聚类和密度估计任务中表现准确,自动确定子集数量。
  • 适合无先验知识的多峰数据建模,如生物统计、金融分析。

单峰性是数据围绕单一密度极值聚集的关键特征。本文提出一种方法,通过递归分割数据密度中的谷点,将单变量数据划分为多个单峰子集。为检测谷点,引入经验累积分布函数(ecdf)凸包上临界点的性质,可指示密度谷的存在。随后,对每个单峰子集应用统一混合模型(UMM)进行统计建模,最终构建初始数据集的分层混合模型——单峰混合模型(UDMM)。该方法为非参数、无超参数,能自动估计单峰子集数量,并在聚类与密度估计任务中表现出高精度。

原文摘要 · Abstract (English)

Unimodality constitutes a key property indicating grouping behavior of the data around a single mode of its density. We propose a method that partitions univariate data into unimodal subsets through recursive splitting around valley points of the data density. For valley point detection, we introduce properties of critical points on the convex hull of the empirical cumulative density function (ecdf) plot that provide indications on the existence of density valleys. Next, we apply a unimodal data modeling approach that provides a statistical model for each obtained unimodal subset in the form of a Uniform Mixture Model (UMM). Consequently, a hierarchical statistical model of the initial dataset is obtained in the form of a mixture of UMMs, named as the Unimodal Mixture Model (UDMM). The proposed method is non-parametric, hyperparameter-free, automatically estimates the number of unimodal subsets and provides accurate statistical models as indicated by experimental results on clustering and density estimation tasks.

多峰建模非参数密度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。