arXiv:2502.18469cs.IR2025-02被引 4

用大模型自动生成更准确的文本主题标签

Using LLM-Based Approaches to Enhance and Automate Topic Labeling

  • 用BERTopic提取主题,再选关键词和摘要输入大模型生成标签
  • 不同选词策略影响标签质量,多样性策略表现更优
  • 提出新评估指标,量化标签对主题文档的语义代表性

主题建模已成为分析文本数据的关键方法,尤其适用于从大规模文档中提取有意义的洞察。然而,传统模型输出多为关键词列表,需人工解读才能精准标注。本研究探索利用大语言模型(LLM)自动化并提升主题标签质量,通过BERTopic进行主题建模后,采用不同策略选取各主题内的关键词与文档摘要,输入LLM生成更具语义合理性的标签。不同策略侧重主导主题或多样性,用于评估其对标签质量的影响。此外,针对主题标签缺乏量化评估手段的问题,提出一种新指标,衡量标签对主题内所有文档的语义代表性。

原文摘要 · Abstract (English)

Topic modeling has become a crucial method for analyzing text data, particularly for extracting meaningful insights from large collections of documents. However, the output of these models typically consists of lists of keywords that require manual interpretation for precise labeling. This study explores the use of Large Language Models (LLMs) to automate and enhance topic labeling by generating more meaningful and contextually appropriate labels. After applying BERTopic for topic modeling, we explore different approaches to select keywords and document summaries within each topic, which are then fed into an LLM to generate labels. Each approach prioritizes different aspects, such as dominant themes or diversity, to assess their impact on label quality. Additionally, recognizing the lack of quantitative methods for evaluating topic labels, we propose a novel metric that measures how semantically representative a label is of all documents within a topic.

主题建模大模型应用标签生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。