提出新方法实现无序依赖的高效聚类,避免传统顺序影响。
Consistent Amortized Clustering via Generative Flow Networks
- 用生成流网络统一策略与奖励,共享能量函数建模
- 在合成与真实数据上聚类效果优于现有方法
- 保证聚类后验一致性,实现数据顺序无关性
用于摊销概率聚类的神经模型可基于集合结构输入生成聚类标签,避免长时间马尔可夫链运行和显式数据似然计算。现有方法如神经聚类过程按顺序为每个数据点打标签,常导致聚类结果高度依赖数据顺序;而按顺序构建完整聚类的方法则无法提供分配概率。本文提出GFNCP,一种新型摊销聚类框架。GFNCP被建模为生成流网络,采用共享的能量基参数化策略与奖励。我们证明流匹配条件等价于聚类后验在边际化下的一致性,进而意味着顺序不变性。GFNCP在合成与真实数据集上的聚类性能均优于现有方法。
原文摘要 · Abstract (English)
Neural models for amortized probabilistic clustering yield samples of cluster labels given a set-structured input, while avoiding lengthy Markov chain runs and the need for explicit data likelihoods. Existing methods which label each data point sequentially, like the Neural Clustering Process, often lead to cluster assignments highly dependent on the data order. Alternatively, methods that sequentially create full clusters, do not provide assignment probabilities. In this paper, we introduce GFNCP, a novel framework for amortized clustering. GFNCP is formulated as a Generative Flow Network with a shared energy-based parametrization of policy and reward. We show that the flow matching conditions are equivalent to consistency of the clustering posterior under marginalization, which in turn implies order invariance. GFNCP also outperforms existing methods in clustering performance on both synthetic and real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。