arXiv:2507.16695cs.CLcs.AI2025-07中稿 · and published at C…被引 3

用可解释的矩阵分解挖掘文本隐含主题并学习词向量

Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM

  • 基于行随机化DEDICOM算法分析文本共现矩阵
  • 同时识别词汇隐含主题并生成可解释词向量
  • 适合关注模型可解释性的自然语言处理研究者

DEDICOM算法为对称与非对称方阵提供独特可解释的矩阵分解方法。本文在文本语料的互信息矩阵上采用新的行随机化DEDICOM变体,以识别词汇中的潜在主题簇,并同步学习可解释的词嵌入。提出一种高效训练约束型DEDICOM的方法,并对其主题建模与词向量性能进行了定性评估。

原文摘要 · Abstract (English)

The DEDICOM algorithm provides a uniquely interpretable matrix factorization method for symmetric and asymmetric square matrices. We employ a new row-stochastic variation of DEDICOM on the pointwise mutual information matrices of text corpora to identify latent topic clusters within the vocabulary and simultaneously learn interpretable word embeddings. We introduce a method to efficiently train a constrained DEDICOM algorithm and a qualitative evaluation of its topic modeling and word embedding performance.

主题建模词向量可解释性矩阵分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。