arXiv:2604.03936stat.MLcs.LG2026-04

改进高维数据的双凸块聚类,自动筛选并加权重要特征。

Biconvex Biclustering

  • 通过双凸优化联合学习特征权重与发现块结构
  • 在模拟和淋巴瘤数据中均优于现有方法,准确识别块结构
  • 适合基因表达等高维数据的聚类分析,可解释性强

本文提出一种双凸修正的凸块聚类方法,以提升其在高维场景下的性能。与预先剔除噪声特征的启发式方法不同,该方法联合学习并动态加权信息特征,同时发现块聚类结构。该方法具有数据自适应性,配备基于近端交替最小化的高效算法,包含超参数调优指导及优化子问题的快速求解方案。理论方面,我们在亚高斯误差下建立了目标函数的有限样本界,并将这些保证推广至输入相似性非均匀的情况。大量模拟结果显示,该方法能稳定恢复潜在块结构,合理分配特征权重,优于对比方法。在淋巴瘤基因微阵列数据上的应用中,成功恢复了与已知分类一致的块聚类,且通过列分组和拟合权重提供了对mRNA样本的额外解释。

原文摘要 · Abstract (English)

This article proposes a biconvex modification to convex biclustering in order to improve its performance in high-dimensional settings. In contrast to heuristics that discard a subset of noisy features a priori, our method jointly learns and accordingly weighs informative features while discovering biclusters. Moreover, the method is adaptive to the data, and is accompanied by an efficient algorithm based on proximal alternating minimization, complete with detailed guidance on hyperparameter tuning and efficient solutions to optimization subproblems. These contributions are theoretically grounded; we establish finite-sample bounds on the objective function under sub-Gaussian errors, and generalize these guarantees to cases where input affinities need not be uniform. Extensive simulation results reveal our method consistently recovers underlying biclusters while weighing and selecting features appropriately, outperforming peer methods. An application to a gene microarray dataset of lymphoma samples recovers biclusters matching an underlying classification, while giving additional interpretation to the mRNA samples via the column groupings and fitted weights.

块聚类高维数据基因分析优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。