用NMF分析新冠文献主题演变,揭示研究热点变迁。
Exploring Topic Trends in COVID-19 Research Literature using Non-Negative Matrix Factorization
- 通过NMF分解文献词频矩阵,挖掘隐藏主题结构。
- 发现研究主题随时间演进,早期聚焦病毒机制,后期转向治疗与疫苗。
- 方法严谨,适合想追踪疫情研究脉络的科研人员。
本文利用非负矩阵分解(NMF)对新冠开放研究数据集(CORD-19)进行主题建模,以揭示海量新冠研究文献中的潜在主题结构及其演变。NMF将文档-词项矩阵分解为两个非负矩阵,有效表示主题及其在文档中的分布,从而反映文档与主题、主题与词汇之间的关联强度。研究采用一系列严格的预处理步骤标准化文本数据,同时保留短语上下文,并使用词频-逆文档频率(tf-idf)进行特征提取,依据词频与罕见度赋权。为确保模型稳健性,进行了稳定性分析,评估不同主题数下的稳定得分,以确定最优主题数量。通过分析,我们追踪了主题在时间维度上的演化轨迹。结果有助于理解新冠研究的知识结构,为该领域未来研究提供重要参考资源。
原文摘要 · Abstract (English)
In this work, we apply topic modeling using Non-Negative Matrix Factorization (NMF) on the COVID-19 Open Research Dataset (CORD-19) to uncover the underlying thematic structure and its evolution within the extensive body of COVID-19 research literature. NMF factorizes the document-term matrix into two non-negative matrices, effectively representing the topics and their distribution across the documents. This helps us see how strongly documents relate to topics and how topics relate to words. We describe the complete methodology which involves a series of rigorous pre-processing steps to standardize the available text data while preserving the context of phrases, and subsequently feature extraction using the term frequency-inverse document frequency (tf-idf), which assigns weights to words based on their frequency and rarity in the dataset. To ensure the robustness of our topic model, we conduct a stability analysis. This process assesses the stability scores of the NMF topic model for different numbers of topics, enabling us to select the optimal number of topics for our analysis. Through our analysis, we track the evolution of topics over time within the CORD-19 dataset. Our findings contribute to the understanding of the knowledge structure of the COVID-19 research landscape, providing a valuable resource for future research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。