arXiv:2507.23220cs.CLcs.LG2025-07中稿 · publication in Tra…被引 7

用稀疏自编码器提取语义特征,让主题模型更懂深层概念。

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders

  • 基于稀疏自编码器的可解释特征空间构建主题
  • 在8个数据集上超越传统与神经主题模型的连贯性
  • 支持通过主题向量控制生成内容,适合需要可控生成的研究

传统主题模型虽能发现大规模文本中的潜在主题,但依赖词袋表示,难以捕捉语义抽象特征。尽管部分神经变体使用更丰富的表示,仍受限于以词表描述主题,无法表达复杂概念。本文提出机制化主题模型(MTMs),在稀疏自编码器(SAEs)学习的可解释特征空间上定义主题,使模型能揭示更深层的概念主题,并以表达性强的特征描述呈现。此外,MTMs是唯一支持通过主题引导向量实现可控文本生成的主题模型。为公平评估,提出基于大语言模型的成对比较框架「topic judge」。在八个数据集上,MTMs在连贯性指标上达到或超过传统及神经基线,且被「topic judge」持续偏好,同时支持有效的大型语言模型引导。

原文摘要 · Abstract (English)

Traditional topic models are effective at uncovering latent themes in large text collections. However, due to their reliance on bag-of-words representations, they struggle to capture semantically abstract features. While some neural variants use richer representations, they are similarly constrained by expressing topics as word lists, which limits their ability to articulate complex topics. We introduce Mechanistic Topic Models (MTMs), a class of topic models that operate on interpretable features learned by sparse autoencoders (SAEs). By defining topics over this semantically rich space, MTMs can reveal deeper conceptual themes with expressive feature descriptions. Moreover, uniquely among topic models, MTMs enable controllable text generation using topic steering vectors. To properly evaluate MTM topics against word list approaches, we propose \textit{topic judge}, an LLM-based pairwise comparison evaluation framework. Across eight datasets, MTMs match or exceed traditional and neural baselines on coherence metrics, are consistently preferred by topic judge, and enable effective LLM steering.

主题模型稀疏编码可控生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。