用大模型增强主题建模,让文本分析更懂复杂语境。
Qualitative Insights Tool (QualIT): LLM Enhanced Topic Modeling
- 结合大模型与聚类方法,提升主题语义理解能力。
- 主题一致性达70%,多样性95.5%,优于基准方法。
- 适合需要深度语义分析的研究场景,如人才管理。
主题建模是挖掘大规模文本语料中主题结构的常用技术。然而,传统方法如隐含狄利克雷分布(LDA)难以捕捉复杂叙事所需的细微语义和上下文理解。近年来,BERTopic等方法显著提升了主题一致性,成为新基准。本文提出一种新方法——定性洞察工具(QualIT),将大语言模型(LLMs)与现有基于聚类的主题建模方法结合。该方法利用大模型的深层上下文理解与强大语言生成能力,增强聚类式主题建模过程。我们在大型新闻文章语料上评估该方法,结果表明其在主题一致性和主题多样性上均有显著提升:在20个真实主题上,主题一致性达到70%(对比基线65%和57%),主题多样性达95.5%(对比基线85%和72%)。研究发现,大模型的融合为动态、复杂的文本数据主题建模开辟了新可能,尤其适用于人才管理等研究场景。
原文摘要 · Abstract (English)
Topic modeling is a widely used technique for uncovering thematic structures from large text corpora. However, most topic modeling approaches e.g. Latent Dirichlet Allocation (LDA) struggle to capture nuanced semantics and contextual understanding required to accurately model complex narratives. Recent advancements in this area include methods like BERTopic, which have demonstrated significantly improved topic coherence and thus established a new standard for benchmarking. In this paper, we present a novel approach, the Qualitative Insights Tool (QualIT) that integrates large language models (LLMs) with existing clustering-based topic modeling approaches. Our method leverages the deep contextual understanding and powerful language generation capabilities of LLMs to enrich the topic modeling process using clustering. We evaluate our approach on a large corpus of news articles and demonstrate substantial improvements in topic coherence and topic diversity compared to baseline topic modeling techniques. On the 20 ground-truth topics, our method shows 70% topic coherence (vs 65% & 57% benchmarks) and 95.5% topic diversity (vs 85% & 72% benchmarks). Our findings suggest that the integration of LLMs can unlock new opportunities for topic modeling of dynamic and complex text data, as is common in talent management research contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。