arXiv:2603.10619cs.CL2026-03被引 1

区分主题模型中的相似性与关联性,提升语义理解精度

Disentangling Similarity and Relatedness in Topic Models

  • 用双轴框架区分词的相似性与关联性,构建合成标注数据集
  • 不同主题模型在相似性-关联性空间中位置各异,影响下游任务表现
  • 同时评估两轴可诊断模型语义结构,适配不同应用场景

大预训练语言模型(PLM)的成功推动其融入主题建模。然而,与经典共现模型(如LDA)相比,PLM增强型主题模型不仅性能不同,捕捉的语义结构也存在差异。本文沿两个心理语言学维度——主题关联性(如狗/骨头)和分类相似性(如狗/狼)——形式化这一区别。我们利用大语言模型构建大规模词对合成基准,并在此上训练神经评分器。在多个语料库和模型族中,该评分器将不同主题模型置于相似性-关联性空间的不同位置。两维得分能有效预测下游任务表现:需相似性的任务受益于高相似性主题,需关联性的任务则偏好高关联性主题;过度强调任一维度都会损害另一类任务性能。因此,单一维度并非普遍有益。同时测量两轴,可为评估主题模型语义结构提供一种实用、模型无关的诊断工具。

原文摘要 · Abstract (English)

The recent success of large pre-trained language models (PLMs) has motivated their integration into topic modeling. However, PLM-augmented topic models differ from classical co-occurrence models such as Latent Dirichlet Allocation (LDA) not only in performance, but also in the type of semantic structure they capture. We formalize this distinction along two psycholinguistic axes: thematic relatedness (dog/bone) and taxonomic similarity (dog/wolf). To measure both axes over topic words, we construct a large synthetic benchmark of word pairs using LLM-based annotation and train a neural scorer on it. Across multiple corpora and model families, the scorer places different topic-model families at distinct positions within the joint similarity-relatedness space. The two scores further predict downstream task performance: tasks requiring similarity benefit from similarity-rich topics, whereas tasks requiring relatedness benefit from the converse, and excessive emphasis on either axis degrades performance on tasks aligned with the opposing semantic structure. Neither axis is uniformly beneficial. Measuring both therefore provides a practical, model-agnostic diagnostic for evaluating the semantic structure captured by topic models.

主题建模语义结构大模型双轴分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。