arXiv:2508.06991cs.LG2025-08被引 1

为可解释的Tsetlin机设计了高效特征选择方法。

A Comparative Study of Feature Selection in Tsetlin Machines

  • 用内部规则权重和状态设计新型特征重要性评分器。
  • 新方法在12个数据集上保持高精度,计算成本更低。
  • 适合追求模型可解释性的机器学习研究者。

特征选择对提升模型可解释性、降低复杂度甚至提高准确率至关重要。最近提出的Tsetlin机(TM)具备基于规则的学习能力,但缺乏成熟的特征重要性评估工具。本文评估了一系列特征选择技术在TM上的表现,包括经典滤波法与嵌入式方法,以及为神经网络设计的后处理解释方法(如SHAP、LIME),还提出一类基于TM规则权重和Tsetlin自动机(TA)状态的新嵌入式评分器。在12个数据集上通过移除重训(ROAR)和移除去偏(ROAD)等评估协议,验证其因果影响。结果表明,TM内生评分器不仅性能优异,还能利用规则可解释性揭示特征交互模式;更简单的特定评分器在极低计算开销下实现相近精度保留。本研究首次建立了TM特征选择的全面基准,为发展专用可解释性技术铺平道路。

原文摘要 · Abstract (English)

Feature Selection (FS) is crucial for improving model interpretability, reducing complexity, and sometimes for enhancing accuracy. The recently introduced Tsetlin machine (TM) offers interpretable clause-based learning, but lacks established tools for estimating feature importance. In this paper, we adapt and evaluate a range of FS techniques for TMs, including classical filter and embedded methods as well as post-hoc explanation methods originally developed for neural networks (e.g., SHAP and LIME) and a novel family of embedded scorers derived from TM clause weights and Tsetlin automaton (TA) states. We benchmark all methods across 12 datasets, using evaluation protocols, like Remove and Retrain (ROAR) strategy and Remove and Debias (ROAD), to assess causal impact. Our results show that TM-internal scorers not only perform competitively but also exploit the interpretability of clauses to reveal interacting feature patterns. Simpler TM-specific scorers achieve similar accuracy retention at a fraction of the computational cost. This study establishes the first comprehensive baseline for FS in TM and paves the way for developing specialized TM-specific interpretability techniques.

特征选择可解释性Tsetlin机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。