arXiv:2505.12353cs.LGcs.AI2025-05ICML

为非线性模型设计了重要数据点采样方法,提升训练效率并支持模型解释。

Importance Sampling for Nonlinear Models

  • 引入伴随算子,将线性模型的范数与杠杆率采样推广至非线性场景。
  • 基于新定义的非线性范数和杠杆率采样,可保证对非线性映射的近似精度。
  • 适用于大规模数据训练加速、模型可解释性分析与异常值检测。

尽管基于范数和杠杆率的方法在线性模型中已被广泛研究用于识别重要数据点,但针对非线性模型的类似工具仍严重不足。通过引入非线性映射的伴随算子,本文填补了这一空白,将范数和杠杆率采样推广至非线性设置。我们证明,基于这些广义范数和杠杆率的采样可为底层非线性映射提供近似保证,类似于线性子空间嵌入。作为直接应用,该方法不仅通过在大规模数据集上高效采样降低非线性模型的训练复杂度,还提供了全新的模型可解释性与异常值检测机制。理论分析与多种监督学习场景下的实验结果验证了本工作的有效性。

原文摘要 · Abstract (English)

While norm-based and leverage-score-based methods have been extensively studied for identifying "important" data points in linear models, analogous tools for nonlinear models remain significantly underdeveloped. By introducing the concept of the adjoint operator of a nonlinear map, we address this gap and generalize norm-based and leverage-score-based importance sampling to nonlinear settings. We demonstrate that sampling based on these generalized notions of norm and leverage scores provides approximation guarantees for the underlying nonlinear mapping, similar to linear subspace embeddings. As direct applications, these nonlinear scores not only reduce the computational complexity of training nonlinear models by enabling efficient sampling over large datasets but also offer a novel mechanism for model explainability and outlier detection. Our contributions are supported by both theoretical analyses and experimental results across a variety of supervised learning scenarios.

非线性模型重要性采样模型可解释性数据压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。