arXiv:2409.14429cs.LGcs.AI2024-09中稿 · publication in Bus…被引 31

挑战模型性能与可解释性的权衡,证明可解释模型也能高精度

Challenging the Performance-Interpretability Trade-off: An Evaluation of Interpretable Machine Learning Models

  • 对比14种模型在20个表格数据集上的表现,采用大规模调参与交叉验证
  • 7种广义加性模型(GAMs)达到与黑箱模型相当的准确率
  • 适合关注模型可信度的信息系统、医疗等需透明决策的领域

机器学习正渗透至各个领域以支持数据驱动决策。以往常因性能优势而偏好复杂黑箱模型,而可解释模型则被认为预测能力较弱。然而,近年来一类新型广义加性模型(GAMs)被提出,可在保持完全可解释性的同时捕捉复杂非线性模式。为揭示其优劣,本研究在20个表格基准数据集上,比较了7种GAMs与7种常用机器学习模型的预测性能。通过大规模超参数搜索结合交叉验证,共完成68,500次模型运行,确保比较公平可靠。同时定性分析模型输出可视化结果,评估其可解释性。结果表明,对于表格数据,性能与可解释性之间并无严格权衡;真正实现高精度与可解释性兼得。本文还从人机协同视角讨论GAMs在信息系统领域的价值,并为未来研究提供启示。

原文摘要 · Abstract (English)

Machine learning is permeating every conceivable domain to promote data-driven decision support. The focus is often on advanced black-box models due to their assumed performance advantages, whereas interpretable models are often associated with inferior predictive qualities. More recently, however, a new generation of generalized additive models (GAMs) has been proposed that offer promising properties for capturing complex, non-linear patterns while remaining fully interpretable. To uncover the merits and limitations of these models, this study examines the predictive performance of seven different GAMs in comparison to seven commonly used machine learning models based on a collection of twenty tabular benchmark datasets. To ensure a fair and robust model comparison, an extensive hyperparameter search combined with cross-validation was performed, resulting in 68,500 model runs. In addition, this study qualitatively examines the visual output of the models to assess their level of interpretability. Based on these results, the paper dispels the misconception that only black-box models can achieve high accuracy by demonstrating that there is no strict trade-off between predictive performance and model interpretability for tabular data. Furthermore, the paper discusses the importance of GAMs as powerful interpretable models for the field of information systems and derives implications for future work from a socio-technical perspective.

可解释性表格数据广义加性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。