arXiv:2603.25958cs.LG2026-03

提出自适应特征提取方法,通过加权k均值优化特征重要性。

Cluster-Adaptive Feature Extraction and its Theoretical Foundation with Minkowski Weighted k-Means

  • 基于Minkowski加权k均值,用指数p控制特征选择强度。
  • 理论证明特征权重与离散度呈幂律关系,可抑制噪声特征。
  • 新方法CAFE在多种数据上提升传统特征提取效果。

Minkowski加权k均值(mwk-means)算法通过引入特征权重和Minkowski距离扩展了经典k均值。我们首次表明,mwk-means目标函数可表示为簇内离散度的幂平均,其阶数由Minkowski指数p决定。该表达揭示了p如何调控特征的选型与均匀使用。基于此,我们推导出目标函数的界,并刻画了特征权重结构:仅依赖相对离散度,且与离散度比值呈幂律关系。由此获得对高离散度特征抑制的明确保证,并建立了算法收敛性。基于这些理论结果,我们提出簇自适应特征提取(CAFE)方法,利用mwk-means特征权重对数据进行重缩放,以实现簇内离散度排序反转,从而抑制噪声特征、增强有用特征。在受控簇内噪声环境下进行的大量实验表明,CAFE持续提升了传统特征提取方法的性能。

原文摘要 · Abstract (English)

The Minkowski weighted $k$-means ($mwk$-means) algorithm extends classical $k$-means by incorporating feature weights and a Minkowski distance. We first show that the $mwk$-means objective can be expressed as a power-mean aggregation of within-cluster dispersions, with the order determined by the Minkowski exponent $p$. This formulation reveals how $p$ controls the transition between selective and uniform use of features. Using this representation, we derive bounds for the objective function and characterise the structure of the feature weights, showing that they depend only on relative dispersion and follow a power-law relationship with dispersion ratios. This leads to explicit guarantees on the suppression of high-dispersion features, and we establish convergence of the algorithm. Building on these theoretical results, we introduce Cluster-Adaptive Feature Extraction (CAFE), a method that uses the $mwk$-means feature weights to rescale the data prior to unsupervised feature extraction. We prove that this rescaling reverses the within-cluster dispersion ordering, suppressing noisy features and amplifying informative ones. Numerous experiments conducted under controlled within-cluster noise show that CAFE consistently improves the results of traditional feature extraction methods.

聚类特征提取加权聚类理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。