让主题模型认识实体,提升对新闻事件演化的捕捉能力。
Embedded Topic Models Enhanced by Wikification
- 引入维基知识增强神经主题模型,识别命名实体
- 在纽约时报和AIDA-CoNLL数据集上提升泛化性能
- 能有效捕捉主题随时间演变的动态过程
主题模型用于分析文档集合以发现词的有意义模式。然而,以往的主题模型仅关注词的拼写,未考虑词的同形异义问题。本研究将维基百科知识融入神经主题模型,使其具备命名实体感知能力。我们在两个数据集上评估方法:1)《纽约时报》新闻文章;2)AIDA-CoNLL数据集。实验表明,该方法显著提升了神经主题模型的泛化能力。此外,通过分析各主题中的高频词及主题间的时序依赖关系,验证了该实体感知模型能够良好捕捉主题的时间序列发展过程。
原文摘要 · Abstract (English)
Topic modeling analyzes a collection of documents to learn meaningful patterns of words. However, previous topic models consider only the spelling of words and do not take into consideration the homography of words. In this study, we incorporate the Wikipedia knowledge into a neural topic model to make it aware of named entities. We evaluate our method on two datasets, 1) news articles of \textit{New York Times} and 2) the AIDA-CoNLL dataset. Our experiments show that our method improves the performance of neural topic models in generalizability. Moreover, we analyze frequent terms in each topic and the temporal dependencies between topics to demonstrate that our entity-aware topic models can capture the time-series development of topics well.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。