基于属性重要性的分类器,提升可解释性与效率
CNC-TP: Classifier Nominal Concept Based on Top-Pertinent Attributes
- 用重要属性构建部分概念格,聚焦关键知识
- 在真实数据集上验证方法,效率优于传统方法
- 适合需要可解释模型的决策场景
知识发现旨在从海量数据中提取隐藏且有意义的知识,涵盖数据选择、预处理、转换、数据挖掘和可视化等步骤。其中分类是核心任务之一,需利用标注数据训练分类器预测新样本类别。现有方法包括决策树、贝叶斯分类器、最近邻、神经网络、支持向量机及形式概念分析(FCA)。FCA因其基于概念格的数学结构,具备良好的可解释性。本文综述了基于FCA的分类器研究进展,探讨从名义数据中计算闭包算子的方法,并提出一种聚焦最相关概念的局部概念格构造新方法。实验结果表明该方法具有较高的效率。
原文摘要 · Abstract (English)
Knowledge Discovery in Databases (KDD) aims to exploit the vast amounts of data generated daily across various domains of computer applications. Its objective is to extract hidden and meaningful knowledge from datasets through a structured process comprising several key steps: data selection, preprocessing, transformation, data mining, and visualization. Among the core data mining techniques are classification and clustering. Classification involves predicting the class of new instances using a classifier trained on labeled data. Several approaches have been proposed in the literature, including Decision Tree Induction, Bayesian classifiers, Nearest Neighbor search, Neural Networks, Support Vector Machines, and Formal Concept Analysis (FCA). The last one is recognized as an effective approach for interpretable and explainable learning. It is grounded in the mathematical structure of the concept lattice, which enables the generation of formal concepts and the discovery of hidden relationships among them. In this paper, we present a state-of-theart review of FCA-based classifiers. We explore various methods for computing closure operators from nominal data and introduce a novel approach for constructing a partial concept lattice that focuses on the most relevant concepts. Experimental results are provided to demonstrate the efficiency of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。