arXiv:2511.12850cs.CL2025-11

LDA模型重复稳定不代表真找到了正确主题,需分开评估稳定性与恢复能力。

Repeatability is not recovery: Quantifying algorithmic stability and topic recovery in Latent Dirichlet Allocation

  • 用合成语料构建真实主题结构,量化模型恢复效果
  • 50次运行中模型稳定但常未还原真实主题,重复性≠正确性
  • 提醒研究者在重要应用中需多维度验证模型输出

主题模型常以多次运行结果的一致性作为优劣判断标准,隐含假设是重复性等于对真实主题的正确恢复。我们证明该假设错误:重复性不等于恢复性。提出一个联合衡量重复性与相对于已知真实结构准确性的稳定性框架。由于真实语料缺乏已知主题结构,我们使用隐狄利克雷分配(LDA)生成过程构造合成语料,实现对主题恢复的直接评估。在每个语料上进行50次重复LDA运行,发现LDA能可靠识别正确主题数量,并频繁收敛到高度一致的主题解。然而,这些重复性高的解往往未能恢复真实生成的主题。因此,内部稳定性不应被解读为正确性的证据。结果表明,稳定性与恢复性是主题模型的两个独立属性,应分别评估。故而,在支持实质性结论前,主题模型输出应通过多种互补标准验证,尤其在高风险应用场景中。

原文摘要 · Abstract (English)

Topic models are often judged by the consistency of their outputs across repeated runs, implicitly assuming that repeatable topic output is a successful recovery of the underlying topics. We show that this assumption is false: repeatability is not recovery. We introduce a stability framework that jointly measures consistency among repeated runs and accuracy relative to known ground truth. Because real-world corpora lack known topic structures, we generate synthetic corpora using the Latent Dirichlet Allocation (LDA) generative process, enabling direct evaluation of topic recovery. Across 50 repeated LDA runs on each corpus, we find that LDA reliably identifies the correct number of topics and frequently converges to highly consistent topic solutions. However, these repeatable solutions frequently fail to recover the true generating topics. Thus, internal stability should not be interpreted as evidence of correctness. Our results illustrate that stability and recovery are distinct properties of topic models and should be evaluated separately. Consequently, topic-model outputs should be validated using multiple complementary criteria before supporting substantive conclusions, particularly in high-stakes applications.

主题模型稳定性真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。