让神经主题模型更懂标签和作者,提升可解释性。
Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors
- 引入元数据增强神经主题模型,融合标签与作者信息。
- 在主题质量和对齐度上优于传统模型,且能识别作者兴趣。
- 适合需要可解释主题分析的研究者与内容平台。
主题模型是提取大规模文档语义信息的常用方法。然而,近期研究表明,这些模型生成的主题常与人类意图不符。尽管标签和作者信息等元数据存在,却未被有效融入神经主题模型。为此,我们提出FANToM,一种新方法,使神经主题模型同时对齐标签与作者信息。该方法在有元数据时可纳入,生成更具可解释性的主题及对应作者分布。实验表明,FANToM通过学习标签、主题与作者间的对齐关系,表达能力更强,显著提升主题质量与对齐度,并能识别作者兴趣与相似性。
原文摘要 · Abstract (English)
Topic models are a popular approach for extracting semantic information from large document collections. However, recent studies suggest that the topics generated by these models often do not align well with human intentions. Although metadata such as labels and authorship information are available, it has not yet been effectively incorporated into neural topic models. To address this gap, we introduce FANToM, a novel method to align neural topic models with both labels and authorship information. FANToM allows for the inclusion of this metadata when available, producing interpretable topics and author distributions for each topic. Our approach demonstrates greater expressiveness than conventional topic models by learning the alignment between labels, topics, and authors. Experimental results show that FANToM improves existing models in terms of both topic quality and alignment. Additionally, it identifies author interests and similarities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。