arXiv:2607.20084stat.MLcs.LG2026-07

对比三个R包在真实数据上的NMF性能,帮研究者选合适工具。

Non--negative matrix factorization using the \textit{R} package \textsf{nnmf}

  • 用真实数据评估NMF包的计算效率和稳定性
  • 新包在收敛性和内存使用上表现更优
  • 适合生物信息、推荐系统等领域的实操用户

非负矩阵分解(NMF)已成为从非负数据中提取潜在结构的重要降维方法,广泛应用于生物信息学、文本挖掘、图像分析和推荐系统等领域。随着NMF流行度上升,多个R包相继推出,采用不同优化策略与计算框架。然而,这些实现方案在真实数据条件下的系统性比较仍较缺乏,导致研究者在实际应用中难以客观选择合适工具。本研究提出一个全新的R包,并与两个主流NMF包进行系统性对比。评估基于真实世界数据,而非模拟数据,以更好反映实际分析中遇到的复杂性、异质性和噪声特征。所有包在统一实验框架下测试,重点考察计算效率、收敛行为、重构精度、内存占用及因子分解结果的稳定性。

原文摘要 · Abstract (English)

Non--negative matrix factorization (NMF) has become an established dimensionality reduction technique for extracting latent structures from non--negative data and has found widespread applications in fields such as bioinformatics, text mining, image analysis, and recommender systems. As the popularity of NMF has increased, numerous \textit{R} packages implementing different optimization strategies and computational frameworks have been developed. Despite their widespread availability, comprehensive evaluations of these implementations under real--world data conditions remain limited. Consequently, researchers often lack objective guidance when selecting an appropriate package for practical applications. This study introduces a new \textit{R} package for NMF and offers asystematic performance comparison with two widely available \textit{R} packages for NMF analysis. Rather than relying on simulated datasets, the evaluation is conducted using real--world data to better reflect the complexity, heterogeneity, and noise characteristics encountered in practical analytical settings. The packages are assessed using a consistent experimental framework, with emphasis on computational efficiency, convergence behavior, reconstruction accuracy, memory utilization, and the stability of the resulting matrix factorization.

NMFR语言降维数据分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。