用大模型实现细粒度主题建模,助力企业文本分析
Industry-Aligned Granular Topic Modeling
- 基于大语言模型构建细粒度主题建模方法
- 在多个真实业务数据集上优于现有主流方法
- 支持文档摘要、主题层级、模型蒸馏等实用功能
主题建模在多个工业领域的文本挖掘与数据分析中应用广泛。尽管粒度(granularity)对业务洞察具有重要意义,但现有主题建模方法生成细粒度主题的能力尚未被充分探索。本文提出TIDE框架,核心为基于大语言模型(LLMs)的新型细粒度主题建模方法,并集成文档摘要、主题层级关系构建和模型蒸馏等辅助功能,以支持实际商业场景。在多种公开数据集与真实企业数据集上的大量实验表明,TIDE的主题建模性能超越现代主流方法,其附加组件有效应对工业级应用场景需求。TIDE框架正在开源过程中。
原文摘要 · Abstract (English)
Topic modeling has extensive applications in text mining and data analysis across various industrial sectors. Although the concept of granularity holds significant value for business applications by providing deeper insights, the capability of topic modeling methods to produce granular topics has not been thoroughly explored. In this context, this paper introduces a framework called TIDE, which primarily provides a novel granular topic modeling method based on large language models (LLMs) as a core feature, along with other useful functionalities for business applications, such as summarizing long documents, topic parenting, and distillation. Through extensive experiments on a variety of public and real-world business datasets, we demonstrate that TIDE's topic modeling approach outperforms modern topic modeling methods, and our auxiliary components provide valuable support for dealing with industrial business scenarios. The TIDE framework is currently undergoing the process of being open sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。