同时稀疏模型与数据,提升回归模型的鲁棒性。
Joint Model and Data Sparsification via the Marginal Likelihood

- 通过联合优化特征与样本重要性,统一实现模型与数据稀疏化。
- 在多种回归任务中均获得更稀疏且抗噪声的预测模型。
- 方法保持共轭性,适合高维数据与异常值场景。
线性系统中的稀疏恢复广泛应用于信号处理与高维回归。稀疏贝叶斯学习基于自动相关性确定(ARD)原则,通过边缘似然优化实现特征稀疏性。然而,其依赖同方差噪声模型,对异常值或噪声误设敏感,影响模型拟合与预测性能。为此,我们提出联合学习特征与样本相关性的方法,通过单一贝叶斯目标实现模型与数据的同时稀疏化。该对称剪枝机制自然扩展了传统方法,保持共轭性,支持闭式更新,符合稳健回归与影响函数的视角。在多样化的回归任务中,实验验证联合ARD方法始终生成稀疏且鲁棒的预测模型。
原文摘要 · Abstract (English)
Sparse recovery in linear systems underpins applications from signal processing to high-dimensional regression. Sparse Bayesian Learning, grounded in the principle of automatic relevance determination (ARD), offers a practical Bayesian mechanism for feature sparsity via marginal likelihood optimization. Yet, its reliance on a homoscedastic noise model renders it sensitive to data contaminations such as outliers or misspecified noise, harming model fit and predictions. Instead, we propose jointly learning individual feature and sample relevancies, enabling simultaneous model and data sparsification via a single Bayesian objective. This symmetric pruning of model and data offers a natural extension that preserves conjugacy, admits closed-form updates for standard optimization procedures, and aligns with perspectives from robust regression and influence functions. Empirical results across diverse regression tasks affirm that a joint ARD approach consistently yields both sparse and robust prediction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。