arXiv:2501.10137cs.HCcs.LG2025-01被引 4

为关键词过滤提供可交互的概率分析,提升主题模型可信度。

Visual Exploration of Stopword Probabilities in Topic Models

  • 基于语料库概率估计关键词出现可能性
  • 用户信心提升,结果更可信且可解释
  • 适合需增强可视化可信度的研究者

停用词移除是许多机器学习方法中的关键步骤,但常被忽视,影响模型可视化并削弱使用者信心。不当或仓促移除停用词不仅导致性能下降,还显著降低模型质量,使从业者和利益相关方难以信赖输出结果。本文提出一种新方法,提供针对特定语料库的停用词概率估计,并设计交互式可视化系统支持分析。我们在真实数据上评估了该方法与接口,结合常用机器学习方法(主题建模)及全面的定性实验,探究用户信心变化。结果显示,本系统通过(1)返回合理概率、(2)生成合适且具代表性的停用词扩展列表、(3)提供可调节阈值实现可视化分析,有效提升了用户对主题模型可信度的信心。最后,我们总结洞察、建议与最佳实践,以支持从业者改进机器学习输出与主题模型可视化效果。

原文摘要 · Abstract (English)

Stopword removal is a critical stage in many Machine Learning methods but often receives little consideration, it interferes with the model visualizations and disrupts user confidence. Inappropriately chosen or hastily omitted stopwords not only lead to suboptimal performance but also significantly affect the quality of models, thus reducing the willingness of practitioners and stakeholders to rely on the output visualizations. This paper proposes a novel extraction method that provides a corpus-specific probabilistic estimation of stopword likelihood and an interactive visualization system to support their analysis. We evaluated our approach and interface using real-world data, a commonly used Machine Learning method (Topic Modelling), and a comprehensive qualitative experiment probing user confidence. The results of our work show that our system increases user confidence in the credibility of topic models by (1) returning reasonable probabilities, (2) generating an appropriate and representative extension of common stopword lists, and (3) providing an adjustable threshold for estimating and analyzing stopwords visually. Finally, we discuss insights, recommendations, and best practices to support practitioners while improving the output of Machine Learning methods and topic model visualizations with robust stopword analysis and removal.

主题模型可视化停用词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。