传统主题模型评估可能失效,需结合大模型语义相似度重评动态主题。
When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics

- 用大模型计算语义相似度,替代依赖词表的旧评估方法。
- 在纽约时报等数据集上,大模型与人工判断相关性达0.721,传统方法仅0.614。
- 建议按词汇变化程度分层评估,避免整体评价掩盖差异。
动态主题模型捕捉词汇分布随时间演变,但传统一致性指标在词汇更替而语义保持时可能失准。我们对CoNTM和DLDA在纽约时报、DBLP及arXiv上的120个主题进行评估,由三位人工标注者划分低、中、高词汇变化等级。传统时间一致性与人工判断的相关性波动较大(ρ=-0.256至0.614)。相比之下,基于大模型的语义相似度在CoNTM的纽约时报(ρ=0.609)、DBLP(ρ=0.721)和arXiv(ρ=0.502)上与人工语义判断高度一致,但对DLDA效果较差。按词汇变化层级分析揭示了聚合评估所掩盖的差异。因此,建议采用词汇变化感知的评估方式,同时报告传统一致性与大模型语义度量,二者作为互补而非互换信号。
原文摘要 · Abstract (English)
Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($ρ$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($ρ$=0.609), DBLP ($ρ$=0.721), and arXiv ($ρ$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。