提出特征有用性新概念,高效判断模型中关键特征。
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
- 基于逻辑推理定义特征相关性与必要性,改进检测算法。
- 证明神经网络中必要特征可高效识别,支持复杂模型。
- 引入全局有用性概念,适用于决策树等多类模型。
针对分类模型的预测结果,现有方法常通过启发式策略对特征重要性进行排序。本文基于命题逻辑与充分理由概念,提出相关性和必要性特征的判定框架,并改进了相应的算法。特别地,证明在神经网络等复杂模型中,必要特征可高效检测。同时,拓展了相关性的定义并研究其关联问题。此外,提出一个新的全局性概念——有用性(usefulness),旨在衡量特征对模型整体行为的重要性,而非仅针对特定输入。证明有用性与相关性、必要性存在理论关联。开发了在决策树及其他复杂模型中检测有用性的高效算法,并在三个数据集上验证其实际效用。
原文摘要 · Abstract (English)
Given a classification model and a prediction for some input, there are heuristic strategies for ranking features according to their importance in regard to the prediction. One common approach to this task is rooted in propositional logic and the notion of \textit{sufficient reason}. Through this concept, the categories of relevant and necessary features were proposed in order to identify the crucial aspects of the input. This paper improves the existing techniques and algorithms for deciding which are the relevant and/or necessary features, showing in particular that necessity can be detected efficiently in complex models such as neural networks. We also generalize the notion of relevancy and study associated problems. Moreover, we present a new global notion (i.e. that intends to explain whether a feature is important for the behavior of the model in general, not depending on a particular input) of \textit{usefulness} and prove that it is related to relevancy and necessity. Furthermore, we develop efficient algorithms for detecting it in decision trees and other more complex models, and experiment on three datasets to analyze its practical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。