AutoTM 2.0优化主题模型,提升分析效率与可解释性。
AutoTM 2.0: Automatic Topic Modeling Framework for Documents Analysis
- 采用新优化流程与分布式计算,加速主题建模。
- 在5个数据集上表现优于旧版,跨语言适用。
- 支持自定义算法与评测指标,适合研究与应用。
本文提出 AutoTM 2.0 框架,用于优化加性正则化主题模型。相比前一版本,新增优化流水线、基于大模型的评估指标和分布式模式。该框架为专家与非专家提供便捷工具,适用于文本探索性分析或可解释特征聚类任务。质量评估采用专设指标(如一致性)及 GPT-4 驱动方法。研究人员可轻松集成新优化算法并适配新评估指标,以提升建模效果并拓展实验。我们在5个具有不同特征的数据集及两种语言上验证,结果表明 AutoTM 2.0 性能优于旧版。
原文摘要 · Abstract (English)
In this work, we present an AutoTM 2.0 framework for optimizing additively regularized topic models. Comparing to the previous version, this version includes such valuable improvements as novel optimization pipeline, LLM-based quality metrics and distributed mode. AutoTM 2.0 is a comfort tool for specialists as well as non-specialists to work with text documents to conduct exploratory data analysis or to perform clustering task on interpretable set of features. Quality evaluation is based on specially developed metrics such as coherence and gpt-4-based approaches. Researchers and practitioners can easily integrate new optimization algorithms and adapt novel metrics to enhance modeling quality and extend their experiments. We show that AutoTM 2.0 achieves better performance compared to the previous AutoTM by providing results on 5 datasets with different features and in two different languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。