用相似性选关键样本,让模型决策像人一样可解释。
Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning

- 基于相似性构建分层选样模型,自动挑选代表性样本
- 在多模态数据上表现优于现有选样方法,且样本数极少
- 适合需要可解释性的金融、医疗等场景
随着决策过程日益复杂,机器学习已成为应对商业与社会挑战的关键工具。然而,许多现有方法的决策过程难以理解。人类常通过将新情况与少数代表性案例对比来做出判断,因此我们设计了一种基于相似性的可解释选样方法,自动选出能代表整体的典型样本(即‘枢轴’)。受决策树启发,提出一种分层、可解释的枢轴选择模型,通过计算输入实例与枢轴之间的相似度进行筛选。该方法既可用于选样,也可独立作为预测模型。进一步拓展至成对枢轴,结合邻近树与斜向树结构,并引入集成学习,显著提升模型灵活性与性能。该方法不依赖特定数据模态,可借助预训练网络处理文本、图像、时间序列和表格数据。在多种数据集上的实验表明,本方法在保持极少量枢轴的前提下,优于其他实例选择策略,并达到与前沿可解释模型相当的精度。
原文摘要 · Abstract (English)
As decision-making processes grow more complex, machine learning tools have become essential for tackling business and societal challenges. However, many existing methods rely on decision-making procedures that are difficult to interpret. Since humans naturally make decisions by comparing new cases with a few representative examples, we aim to design an approach that selects such pivots to construct an interpretable predictive model. Inspired by decision trees, we propose a hierarchical, interpretable-by-design pivot selection model based on the similarity between pivots and input instances. Our method functions both as a pivot selection technique and a standalone predictive model. Extending beyond single pivots, we incorporate pairs of pivots that are used by proximity and oblique trees, as well as ensembles, which enhance the versatility and effectiveness of our proposal. Additionally, our approach is data modality-agnostic, leveraging pre-trained networks for data transformation. Experiments across diverse datasets, including tabular data, text, images, and time series, demonstrate the effectiveness of our approach, outperforming alternative instance selection strategies and achieving competitive results against state-of-the-art interpretable models while maintaining a minimal number of pivots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。