arXiv:2604.14810stat.MLcs.LG2026-04中稿 · AISTATS 2026

用分块策略让SMC算法高效处理大规模文本聚类问题

Scalable Model-Based Clustering with Sequential Monte Carlo

论文配图:Scalable Model-Based Clustering with Sequential Monte Carlo
图 1 · 摘自论文原文
  • 将聚类问题分解为近似独立子问题,压缩状态表示
  • 在真实知识库构建场景中实现高精度与高效率
  • 适合处理复杂分布的大规模在线聚类任务

在线聚类问题中,簇分配存在大量不确定性,需更多数据才能澄清。当簇遵循复杂分布(如文本数据)时,这一问题更难解决。序列蒙特卡洛(SMC)方法能自然地随时间表示和更新不确定性,但对大规模问题内存开销过大。本文提出一种新型SMC算法,通过将聚类问题分解为近似独立的子问题,实现更紧凑的算法状态表示。该方法受知识库构建任务启发,在此类场景及其他传统SMC难以应对的问题中均表现出准确且高效的性能。

原文摘要 · Abstract (English)

In online clustering problems, there is often a large amount of uncertainty over possible cluster assignments that cannot be resolved until more data are observed. This difficulty is compounded when clusters follow complex distributions, as is the case with text data. Sequential Monte Carlo (SMC) methods give a natural way of representing and updating this uncertainty over time, but have prohibitive memory requirements for large-scale problems. We propose a novel SMC algorithm that decomposes clustering problems into approximately independent subproblems, allowing a more compact representation of the algorithm state. Our approach is motivated by the knowledge base construction problem, and we show that our method is able to accurately and efficiently solve clustering problems in this setting and others where traditional SMC struggles.

聚类SMC在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。