arXiv:2504.04079cs.LGstat.ML2025-04被引 1

提出首个端到端可扩展的鲁棒贝叶斯共聚类框架,有效处理高维噪声数据。

Scalable Robust Bayesian Co-Clustering with Compositional ELBOs

  • 采用双重重参数化变分下界,提升梯度信号质量。
  • 在多模态真实数据上,准确率与鲁棒性均优于现有方法。
  • 适合处理高维、缺失或含噪数据,尤其适用于多模态分析场景。

共聚类利用实例与特征之间的对偶性,同时发现两个维度上的有意义分组,在高维或稀疏数据中常优于传统聚类。尽管近期深度学习方法成功融合了特征学习与聚类分配,但仍易受噪声影响,并在标准自编码器中出现后验崩溃。本文首次提出一个完整的变分共聚类框架,直接在隐空间中学习行和列的聚类,利用双重重参数化证据下界(ELBO)改善梯度信噪比。我们的无监督模型结合变分深度嵌入与高斯混合模型(GMM)先验,为实例和特征提供内置聚类机制,自然将隐空间模式对齐至行与列聚类。此外,通过正则化的端到端噪声学习组合式ELBO架构,在联合重构数据的同时,利用KL散度正则化对抗噪声,从而在单一训练流程中优雅处理损坏或缺失输入。为缓解后验崩溃,引入尺度修正:仅在重构路径中增加编码器隐变量均值,保持丰富表示且不扩大KL项。最后,基于互信息的交叉损失确保行与列的协同聚类一致性。在多种模态的真实世界数据集(数值、文本、图像)上的实证结果表明,本方法不仅保留了先前共聚类方法的优势,还在准确率与鲁棒性上更进一步,尤其在高维或噪声环境下表现突出。

原文摘要 · Abstract (English)

Co-clustering exploits the duality of instances and features to simultaneously uncover meaningful groups in both dimensions, often outperforming traditional clustering in high-dimensional or sparse data settings. Although recent deep learning approaches successfully integrate feature learning and cluster assignment, they remain susceptible to noise and can suffer from posterior collapse within standard autoencoders. In this paper, we present the first fully variational Co-clustering framework that directly learns row and column clusters in the latent space, leveraging a doubly reparameterized ELBO to improve gradient signal-to-noise separation. Our unsupervised model integrates a Variational Deep Embedding with a Gaussian Mixture Model (GMM) prior for both instances and features, providing a built-in clustering mechanism that naturally aligns latent modes with row and column clusters. Furthermore, our regularized end-to-end noise learning Compositional ELBO architecture jointly reconstructs the data while regularizing against noise through the KL divergence, thus gracefully handling corrupted or missing inputs in a single training pipeline. To counteract posterior collapse, we introduce a scale modification that increases the encoder's latent means only in the reconstruction pathway, preserving richer latent representations without inflating the KL term. Finally, a mutual information-based cross-loss ensures coherent co-clustering of rows and columns. Empirical results on diverse real-world datasets from multiple modalities, numerical, textual, and image-based, demonstrate that our method not only preserves the advantages of prior Co-clustering approaches but also exceeds them in accuracy and robustness, particularly in high-dimensional or noisy settings.

共聚类贝叶斯方法鲁棒性变分推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。