用NLP自动分类儿童言语障碍文献,发现14个临床相关主题
Technical Report on classification of literature related to children speech disorder
- 结合LDA与BERTopic模型分析4804篇文献摘要,识别潜在主题
- LDA模型主题一致性达0.42,困惑度-7.5,表现稳定;BERTopic异常主题占比低于20%
- 专为语音病理学设计停用词表,提升分类相关性,适合研究者快速定位方向
本技术报告提出一种基于自然语言处理(NLP)的方法,用于系统分类儿童言语障碍领域的科学文献。我们从PubMed数据库中检索并筛选出2015年后发表的4,804篇相关文章,使用领域特定关键词。对摘要进行清洗与预处理后,采用两种主题建模技术——隐狄利克雷分配(LDA)和BERTopic——识别语料库中的潜在主题结构。模型识别出14个具有临床意义的聚类,如婴儿多动与异常癫痫行为。为提高相关性与精度,引入了针对语音病理学定制的停用词列表。评估结果显示,LDA模型的一致性得分为0.42,困惑度为-7.5,表明主题一致性高且预测性能强;BERTopic模型的异常主题比例低于20%,显示出对异构文献的有效分类能力。研究成果为语音语言病理学领域的自动化文献综述提供了基础。
原文摘要 · Abstract (English)
This technical report presents a natural language processing (NLP)-based approach for systematically classifying scientific literature on childhood speech disorders. We retrieved and filtered 4,804 relevant articles published after 2015 from the PubMed database using domain-specific keywords. After cleaning and pre-processing the abstracts, we applied two topic modeling techniques - Latent Dirichlet Allocation (LDA) and BERTopic - to identify latent thematic structures in the corpus. Our models uncovered 14 clinically meaningful clusters, such as infantile hyperactivity and abnormal epileptic behavior. To improve relevance and precision, we incorporated a custom stop word list tailored to speech pathology. Evaluation results showed that the LDA model achieved a coherence score of 0.42 and a perplexity of -7.5, indicating strong topic coherence and predictive performance. The BERTopic model exhibited a low proportion of outlier topics (less than 20%), demonstrating its capacity to classify heterogeneous literature effectively. These results provide a foundation for automating literature reviews in speech-language pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。