对比大模型与概念分析在主题建模中的表现
Large Language Model and Formal Concept Analysis: a comparative study for Topic Modeling
- 用三步提示词零样本生成主题,结合大模型批量处理能力
- 在教学材料和信息系统论文上验证,主题与领域契合度高
- 首次系统比较FCA与LLM在真实任务中的优劣,适合文本分析研究者
主题建模应用日益广泛,从文档检索到情感分析、文本摘要均有涉及。大语言模型(LLM)当前是文本处理主流,但其在主题建模中的有效性研究较少。形式概念分析(FCA)近年被提出作为候选方法,但缺乏实际案例验证。本文比较了LLM与FCA在主题建模中的表现:采用先前用于主题建模与可视化的CREA流程评估FCA,使用GPT-5进行零样本实验,通过三步提示策略实现文档批次的主题生成、结果合并与主题标注。第一项实验复现此前用于评估CREA的教学材料数据集,第二项实验分析40篇信息系统领域的研究论文,将提取主题与潜在子领域进行对比。
原文摘要 · Abstract (English)
Topic modeling is a research field finding increasing applications: historically from document retrieving, to sentiment analysis and text summarization. Large Language Models (LLM) are currently a major trend in text processing, but few works study their usefulness for this task. Formal Concept Analysis (FCA) has recently been presented as a candidate for topic modeling, but no real applied case study has been conducted. In this work, we compare LLM and FCA to better understand their strengths and weakneses in the topic modeling field. FCA is evaluated through the CREA pipeline used in past experiments on topic modeling and visualization, whereas GPT-5 is used for the LLM. A strategy based on three prompts is applied with GPT-5 in a zero-shot setup: topic generation from document batches, merging of batch results into final topics, and topic labeling. A first experiment reuses the teaching materials previously used to evaluate CREA, while a second experiment analyzes 40 research articles in information systems to compare the extracted topics with the underling subfields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。