用KL散度优化非负矩阵分解,提升文档与图像数据的聚类效果。
Orthogonal Nonnegative Matrix Factorization with the Kullback-Leibler divergence
- 基于交替优化的KL散度算法,适用于稀疏计数数据
- 在文档分类和高光谱解混任务中表现优于传统方法
- 适合处理词频、图像像素计数等泊松分布数据
正交非负矩阵分解(ONMF)已成为聚类的标准方法。目前大多数ONMF研究依赖弗罗贝尼乌斯范数评估近似质量。本文提出一种基于KL散度的新模型与算法,称为KL-ONMF。与假设高斯噪声的弗罗贝尼乌斯范数不同,KL散度是泊松分布数据的最大似然估计器,更适合建模文档中的词频向量和图像中的像素计数。我们开发了基于交替优化的算法,并在文档分类和高光谱图像解混任务中验证其性能,结果表明该方法优于基于弗罗贝尼乌斯范数的ONMF。
原文摘要 · Abstract (English)
Orthogonal nonnegative matrix factorization (ONMF) has become a standard approach for clustering. As far as we know, most works on ONMF rely on the Frobenius norm to assess the quality of the approximation. This paper presents a new model and algorithm for ONMF that minimizes the Kullback-Leibler (KL) divergence. As opposed to the Frobenius norm which assumes Gaussian noise, the KL divergence is the maximum likelihood estimator for Poisson-distributed data, which can model better sparse vectors of word counts in document data sets and photo counting processes in imaging. We develop an algorithm based on alternating optimization, KL-ONMF, and show that it performs favorably with the Frobenius-norm based ONMF for document classification and hyperspectral image unmixing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。