为Tsetlin机模型提供概率评分与不确定性量化,提升可解释性。
Uncertainty Quantification in the Tsetlin Machine
- 基于学习动态分析,推导出模型的固有概率评分
- 模拟数据验证概率评分与真实数据概率一致
- 在CIFAR-10图像分类中揭示模型外推时的低置信度
使用Tsetlin机(TMs)进行数据建模本质上是从数据特征中构建逻辑规则,其决策依赖于这些规则的组合,因此模型完全透明,可解释预测过程。本文提出一种针对TM预测的概率评分,并开发新的不确定性量化技术以进一步增强可解释性。该概率评分是任何TM变体的内在属性,通过分析模型学习动态得出。利用模拟数据验证了学习到的TM概率评分与底层数据概率之间存在清晰关联。可视化结果表明,当超出训练数据范围时,模型置信度明显下降,这与人工神经网络常见的外推现象形成对比。最后,将不确定性量化方法应用于基于CIFAR-10数据集的图像分类任务,提供了新见解并指出了当前TM图像分类模型的改进方向。
原文摘要 · Abstract (English)
Data modeling using Tsetlin machines (TMs) is all about building logical rules from the data features. The decisions of the model are based on a combination of these logical rules. Hence, the model is fully transparent and it is possible to get explanations of its predictions. In this paper, we present a probability score for TM predictions and develop new techniques for uncertainty quantification to increase the explainability further. The probability score is an inherent property of any TM variant and is derived through an analysis of the TM learning dynamics. Simulated data is used to show a clear connection between the learned TM probability scores and the underlying probabilities of the data. A visualization of the probability scores also reveals that the TM is less confident in its predictions outside the training data domain, which contrasts the typical extrapolation phenomenon found in Artificial Neural Networks. The paper concludes with an application of the uncertainty quantification techniques on an image classification task using the CIFAR-10 dataset, where they provide new insights and suggest possible improvements to current TM image classification models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。