提出新方法评估主题模型,让人类判断主题间是否清晰区分。
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
- 设计新标注任务TWM,测试人类能否分辨单个或混合主题的词汇
- 在近4000次标注中发现人工判断与自动指标不总一致
- 适合研究领域特定文本的主题模型评估者使用
主题模型可挖掘文本中的潜在主题结构,但其质量评估仍具挑战性,尤其在专业领域。现有方法多依赖自动化指标如主题连贯性与多样性,但可能与人类判断不符。人工评估任务如词插入测试虽有价值,但成本高,且主要验证于通用语料。本文提出主题词混合(TWM)新任务,通过测试标注者能否区分单一主题或混合主题的词集,评估主题间差异性。TWM补充了词插入任务对主题内连贯性的关注,为多样性指标提供人类基准。我们在哲学科学领域的专用语料上,对六种主题模型(LDA、NMF、Top2Vec、BERTopic、CFMF、CFMF-emb)进行了比较,基于近4000次标注结果发现:词插入测试与连贯性指标在专业领域常不一致;而TWM能捕捉人类感知的主题区分度,并与多样性指标呈现良好一致性。论文发布标注数据集及任务生成代码。研究强调需构建融合自动化与人工评估的框架,尤其针对领域特定语料。
原文摘要 · Abstract (English)
Topic models uncover latent thematic structures in text corpora, yet evaluating their quality remains challenging, particularly in specialized domains. Existing methods often rely on automated metrics like topic coherence and diversity, which may not fully align with human judgment. Human evaluation tasks, such as word intrusion, provide valuable insights but are costly and primarily validated on general-domain corpora. This paper introduces Topic Word Mixing (TWM), a novel human evaluation task assessing inter-topic distinctness by testing whether annotators can distinguish between word sets from single or mixed topics. TWM complements word intrusion's focus on intra-topic coherence and provides a human-grounded counterpart to diversity metrics. We evaluate six topic models - both statistical and embedding-based (LDA, NMF, Top2Vec, BERTopic, CFMF, CFMF-emb) - comparing automated metrics with human evaluation methods based on nearly 4,000 annotations from a domain-specific corpus of philosophy of science publications. Our findings reveal that word intrusion and coherence metrics do not always align, particularly in specialized domains, and that TWM captures human-perceived distinctness while appearing to align with diversity metrics. We release the annotated dataset and task generation code. This work highlights the need for evaluation frameworks bridging automated and human assessments, particularly for domain-specific corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。