新模型自动学习数据权重,提升高维数据中复杂噪声下的鲁棒性。
Meta Additive Model: Interpretable Sparse Learning With Auto Weighting

- 用双层优化框架让模型自适应学习样本权重,无需人工设定函数。
- 在含异常值、标签噪声等复杂噪声下,性能优于现有主流加法模型。
- 适合需要可解释性与抗噪能力的高维变量选择和不平衡分类任务。
稀疏加法模型因其灵活表示与强可解释性,在高维数据分析中备受关注。然而,现有模型多基于均方误差准则进行单层学习,面对非高斯扰动、异常值、噪声标签和类别不平衡等复杂噪声时,实证性能显著下降。样本重加权策略虽被广泛采用以降低对异常数据的敏感性,但通常需预先设定权重函数并手动调节额外超参数。为此,本文提出一种基于双层优化框架的元加法模型(MAM),通过在元数据上训练的MLP参数化权重函数,实现数据驱动的个体损失权重学习。MAM可支持变量选择、鲁棒回归估计与不平衡分类等多种任务。理论上,在温和条件下,MAM保证了计算收敛性、算法泛化性及变量选择一致性。实验表明,无论在合成数据还是真实数据上,面对多种数据污染情形,MAM均显著优于多个前沿加法模型。
原文摘要 · Abstract (English)
Sparse additive models have attracted much attention in high-dimensional data analysis due to their flexible representation and strong interpretability. However, most existing models are limited to single-level learning under the mean-squared error criterion, whose empirical performance can degrade significantly in the presence of complex noise, such as non-Gaussian perturbations, outliers, noisy labels, and imbalanced categories. The sample reweighting strategy is widely used to reduce the model's sensitivity to atypical data; however, it typically requires prespecifying the weighting functions and manually selecting additional hyperparameters. To address this issue, we propose a new meta additive model (MAM) based on the bilevel optimization framework, which learns data-driven weighting of individual losses by parameterizing the weighting function via an MLP trained on meta data. MAM is capable of a variety of learning tasks, including variable selection, robust regression estimation, and imbalanced classification. Theoretically, MAM provides guarantees on convergence in computation, algorithmic generalization, and variable selection consistency under mild conditions. Empirically, MAM outperforms several state-of-the-art additive models on both synthetic and real-world data under various data corruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。