将分层非负矩阵分解扩展到张量,更好保留多模态数据结构。
Stratified Non-Negative Tensor Factorization
- 提出分层非负张量分解(Stratified-NTF),保留多模态数据的几何结构
- 在文本和图像数据上实现更低内存消耗的可解释主题提取
- 适用于多源异构数据,尤其适合关注数据结构与共享模式的研究者
非负矩阵分解(NMF)和非负张量分解(NTF)能将非负高维数据分解为非负低秩成分,因其内在可解释性及对大规模数据的有效性而广受欢迎。近期研究提出分层非负矩阵分解(Stratified-NMF),用于处理来自不同来源(分层)且分布不同的数据,旨在同时恢复分层特异性信息与跨分层共享的全局主题。然而,将Stratified-NMF应用于多模态数据需沿模态展开,导致隐含的张量几何结构丢失。为此,本文将Stratified-NMF推广至张量框架,设计了乘法更新规则,并在文本与图像数据上进行了验证。结果表明,Stratified-NTF能识别出更具可解释性的主题,且内存开销低于Stratified-NMF。此外,我们还引入正则化版本,并在图像数据上展示了其有效性能。
原文摘要 · Abstract (English)
Non-negative matrix factorization (NMF) and non-negative tensor factorization (NTF) decompose non-negative high-dimensional data into non-negative low-rank components. NMF and NTF methods are popular for their intrinsic interpretability and effectiveness on large-scale data. Recent work developed Stratified-NMF, which applies NMF to regimes where data may come from different sources (strata) with different underlying distributions, and seeks to recover both strata-dependent information and global topics shared across strata. Applying Stratified-NMF to multi-modal data requires flattening across modes, and therefore loses geometric structure contained implicitly within the tensor. To address this problem, we extend Stratified-NMF to the tensor setting by developing a multiplicative update rule and demonstrating the method on text and image data. We find that Stratified-NTF can identify interpretable topics with lower memory requirements than Stratified-NMF. We also introduce a regularized version of the method and demonstrate its effects on image data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。