用大模型生成文本,让主题模型更易懂、更对口研究问题。
Creating Targeted, Interpretable Topic Models with LLM-Generated Text Augmentation
- 用GPT-4生成文本增强原始数据,提升主题模型可解释性。
- 在政治科学案例中,生成主题仅需少量人工干预即可回答具体问题。
- 适合需要精准解读的社科研究者,尤其关注主题可读性与领域适配性。
无监督机器学习方法如主题建模和聚类常用于识别政治学、社会学等领域中非结构化文本数据的潜在模式。这些方法缓解了人工质性分析成本高、可复现性差的问题。然而,主题模型存在可解释性差和难以应对特定领域研究问题两大局限。本文探索利用大语言模型生成文本来增强主题建模输出的潜力。以政治科学为例进行评估,结果表明使用GPT-4生成文本增强后,主题模型能产出高度可解释的主题类别,仅需极少人工指导即可用于解答具体的领域研究问题。
原文摘要 · Abstract (English)
Unsupervised machine learning techniques, such as topic modeling and clustering, are often used to identify latent patterns in unstructured text data in fields such as political science and sociology. These methods overcome common concerns about reproducibility and costliness involved in the labor-intensive process of human qualitative analysis. However, two major limitations of topic models are their interpretability and their practicality for answering targeted, domain-specific social science research questions. In this work, we investigate opportunities for using LLM-generated text augmentation to improve the usefulness of topic modeling output. We use a political science case study to evaluate our results in a domain-specific application, and find that topic modeling using GPT-4 augmentations creates highly interpretable categories that can be used to investigate domain-specific research questions with minimal human guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。