通过正交多项式核揭示SVM决策函数的结构可解释性
Structural interpretability in SVMs with truncated orthogonal polynomial kernels

- 基于截断正交多项式核的SVM,可精确展开为希尔伯特空间坐标
- 提出OKC指数量化模型复杂度在不同交互阶、总次数等维度的分布
- 无需重训练即可诊断模型结构,适合关注模型内在机制的研究者
我们研究了基于截断正交多项式核构建的支持向量机(SVM)的后训练可解释性。由于对应的再生核希尔伯特空间(RKHS)是有限维的,并具有显式的张量积正交基,拟合的决策函数可精确展开为内在的RKHS坐标。这引出了基于归一化正交核贡献(OKC)指数的诊断框架——正交表示贡献分析(ORCA)。这些指数量化了分类器平方RKHS范数在交互阶、总多项式次数、边际坐标效应和成对贡献中的分布情况。该方法完全为后训练,无需代理模型或重新训练。我们在一个合成的双螺旋问题和一个真实的五维超声心动图数据集上展示了其诊断价值。结果表明,所提出的指标揭示了仅凭预测准确率无法捕捉的模型复杂性结构特征。
原文摘要 · Abstract (English)
We study post-training interpretability for Support Vector Machines (SVMs) built from truncated orthogonal polynomial kernels. Since the associated reproducing kernel Hilbert space is finite-dimensional and admits an explicit tensor-product orthonormal basis, the fitted decision function can be expanded exactly in intrinsic RKHS coordinates. This leads to Orthogonal Representation Contribution Analysis (ORCA), a diagnostic framework based on normalized Orthogonal Kernel Contribution (OKC) indices. These indices quantify how the squared RKHS norm of the classifier is distributed across interaction orders, total polynomial degrees, marginal coordinate effects, and pairwise contributions. The methodology is fully post-training and requires neither surrogate models nor retraining. We illustrate its diagnostic value on a synthetic double-spiral problem and on a real five-dimensional echocardiogram dataset. The results show that the proposed indices reveal structural aspects of model complexity that are not captured by predictive accuracy alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。