用大模型评估主题模型,发现传统方法忽略的语义问题。
Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models
- 用九个大模型指标从词汇、语义、结构等四方面评估主题模型质量。
- 实验证明大模型能识别冗余和语义漂移,传统指标常遗漏这些缺陷。
- 适合需要精准评估动态数据主题模型的研究者与系统开发者。
本研究提出一种基于大语言模型(LLM)的自动化框架,用于评估动态演化的主题模型。主题建模在数字图书馆中对组织和检索学术内容至关重要,帮助用户应对复杂且不断变化的知识领域。然而,常用的自动评估指标如一致性与多样性仅捕捉狭窄的统计模式,难以解释实际中的语义失败。本文引入一种面向目标的评估框架,采用九个基于LLM的指标,涵盖主题质量的四个关键维度:词汇有效性、主题内语义合理性、主题间结构合理性以及文档-主题对齐合理性。该框架通过对抗性与采样协议验证,并应用于新闻、学术论文及社交媒体帖子等多类数据集,覆盖多种主题建模方法和开源大模型。分析显示,基于大模型的指标可提供可解释、稳健且任务相关的评估结果,揭示了传统指标常忽视的主题模型关键缺陷,如冗余和语义漂移。这些成果支持开发可扩展、细粒度的评估工具,以维持动态数据中主题的相关性。所有代码与数据已公开于 https://github.com/zhiyintan/topic-model-LLMjudgment。
原文摘要 · Abstract (English)
This study presents a framework for automated evaluation of dynamically evolving topic models using Large Language Models (LLMs). Topic modeling is essential for organizing and retrieving scholarly content in digital library systems, helping users navigate complex and evolving knowledge domains. However, widely used automated metrics, such as coherence and diversity, often capture only narrow statistical patterns and fail to explain semantic failures in practice. We introduce a purpose-oriented evaluation framework that employs nine LLM-based metrics spanning four key dimensions of topic quality: lexical validity, intra-topic semantic soundness, inter-topic structural soundness, and document-topic alignment soundness. The framework is validated through adversarial and sampling-based protocols, and is applied across datasets spanning news articles, scholarly publications, and social media posts, as well as multiple topic modeling methods and open-source LLMs. Our analysis shows that LLM-based metrics provide interpretable, robust, and task-relevant assessments, uncovering critical weaknesses in topic models such as redundancy and semantic drift, which are often missed by traditional metrics. These results support the development of scalable, fine-grained evaluation tools for maintaining topic relevance in dynamic datasets. All code and data supporting this work are accessible at https://github.com/zhiyintan/topic-model-LLMjudgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。