用大模型注意力机制重构可解释主题模型,提升长文本建模能力
LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability
- 基于注意力机制设计白盒主题模型,还原文档-主题与主题-词分布
- 黑盒模型通过长输入生成与信号补偿,实现媲美或超越基线的性能
- 适合需要可解释性与长文本分析的研究者使用
主题建模旨在从语料中生成可解释的主题表示及主题-文档对应关系,但传统神经主题模型(NTMs)受限于表达假设和语义抽象能力。本文从白盒与黑盒视角研究大模型(LLM)在主题建模中的应用。针对白盒LLM,提出注意力感知框架,恢复类NTM的可解释结构,包括文档-主题与主题-词分布,验证了LLM可作为注意力感知的NTM。针对黑盒LLM,将主题建模重构为结构化长输入任务,引入基于多样化主题提示与混合检索的后生成信号补偿方法。实验表明,恢复的注意力结构支持有效主题分配与关键词提取,黑盒长上下文LLM性能达到或超过其他基线。结果揭示了LLM与NTMs间的联系,凸显长上下文LLM在主题建模中的潜力。
原文摘要 · Abstract (English)
Topic modeling aims to produce interpretable topic representations and topic--document correspondences from corpora, but classical neural topic models (NTMs) remain constrained by limited representation assumptions and semantic abstraction ability. We study LLM-based topic modeling from both white-box and black-box perspectives. For white-box LLMs, we propose an attention-informed framework that recovers interpretable structures analogous to those in NTMs, including document-topic and topic-word distributions. This validates the view that LLM can serve as an attention-informed NTM. For black-box LLMs, we reformulate topic modeling as a structured long-input task and introduce a post-generation signal compensation method based on diversified topic cues and hybrid retrieval. Experiments show that recovered attention structures support effective topic assignment and keyword extraction, while black-box long-context LLMs achieve competitive or stronger performance than other baselines. These findings suggest a connection between LLMs and NTMs and highlight the promise of long-context LLMs for topic modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。