比较互信息与敏感性分析在客户营销中的特征选择效果。
Mutual information and sensitivity analysis for feature selection in customer targeting: a comparative study

- 用互信息和数据敏感性分析筛选影响营销成功率的特征
- 敏感性分析用9个特征表现更好,尤其在低误报率时
- 互信息适合想降成本又不丢成功率的银行决策者
特征选择是数据驱动知识发现中的关键任务。本文对比了互信息与近年来兴起的数据敏感性分析在客户营销场景中的表现。针对某银行电话营销案例,分别使用两种方法从13个特征中选出对营销成功影响最大的特征集(互信息选13个,敏感性分析选9个),并基于各自选中的特征构建逻辑回归模型。结果表明:当误报率较低时,敏感性分析表现更优;而在高误报率情况下,互信息略胜一筹。因此,若银行希望小幅降低营销成本且不显著损失成功率,互信息仍是优选。研究证实,尽管互信息并非新方法,但在实际应用中仍具有效性;而数据敏感性分析则以更少特征获得良好预测性能。
原文摘要 · Abstract (English)
Feature selection is a highly relevant task in a data-driven knowledge discovery project. Several techniques have been developed aiming at finding the features that influence most an outcome to predict, including mutual information and, in recent years, the data-based sensitivity analysis. The present research focus on analyzing the advantages and disadvantages of each of these two techniques, by applying both to a bank telemarketing case. Thereafter, a logistic regression model is built on the tuned set of features identified by each of the two techniques as the most influencing set of features on the success of a telemarketing contact, in a total of 13 features for mutual information and 9 features for the data-based sensitivity analysis. The latter performs better for lower values of false positives while the former is slightly better for a higher false positive ratio. Thus, mutual information becomes a better choice if bank managers intend to reduce slightly the cost of contacts without risking losing a high number of successes. Such results show that mutual information, although not recent, is still a valid method for feature selection. On the other side, the data-based sensitivity analysis selection achieved good prediction results with less features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。