通过谢尔比空间重构数据,让分类边界图更清晰准确。
ShapDBM: Exploring Decision Boundary Maps in Shapley Space
- 将数据映射到谢尔比空间再降维,避免传统方法的类别混杂问题。
- 生成的决策边界图在质量指标上相当或更优,且区域更紧凑。
- 适合需要可解释性分析的复杂模型开发者使用。
决策边界图(DBMs)是可视化机器学习分类边界的有效工具。然而,其质量高度依赖于降维(DR)技术及原始高维数据空间的选择。对于复杂机器学习数据,传统降维可能导致类别混淆,使生成的DBM难以使用甚至产生误导。本文提出一种新方法:先将数据空间转换至谢尔比空间,再在此空间中进行降维计算DBM。与直接从原始数据计算相比,新方法生成的决策边界图在质量度量值上表现相当或更优,且决策区域更紧凑、更易探索,与实际模型性能更加一致。
原文摘要 · Abstract (English)
Decision Boundary Maps (DBMs) are an effective tool for visualising machine learning classification boundaries. Yet, DBM quality strongly depends on the dimensionality reduction (DR) technique and high dimensional space used for the data points. For complex ML data, DR can create many mixed classes which yield DBMs that are hard to use or even misleading. We propose a new technique to compute DBMs by transforming data space into Shapley space and computing DR on it. Compared to DBMs computed directly from data, our maps have similar or higher quality metric values and visibly more compact, easier to explore, decision zones that better agree with measured model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。