S2MAM自动筛选重要变量,提升半监督学习的鲁棒性与可解释性。
S2MAM: Semi-supervised Meta Additive Model for Robust Estimation and Variable Selection

- 基于双层优化自动选择关键变量并动态更新相似度矩阵
- 在4个合成与12个真实数据集上均表现稳定,抗噪声能力强
- 适合需要可解释性与高鲁棒性的机器学习应用
半监督学习结合流形正则化是利用标记与未标记数据的经典框架,其核心假设为未知边缘分布具有黎曼流形的几何结构。通常通过图拉普拉斯矩阵近似拉普拉斯-贝尔特拉米算子的流形正则化,但该矩阵依赖预设相似度度量,在冗余或噪声输入变量下可能导致不当惩罚。为此,本文提出一种新的半监督元加法模型(S²MAM),基于双层优化机制,自动识别信息变量、更新相似度矩阵,并实现可解释预测。理论分析提供了计算收敛性与统计泛化界保证。在4个合成数据集和12个真实世界数据集上,经受不同水平与类型的污染测试,验证了方法的鲁棒性与可解释性。
原文摘要 · Abstract (English)
Semi-supervised learning with manifold regularization is a classical framework for jointly learning from both labeled and unlabeled data, where the key requirement is that the support of the unknown marginal distribution has the geometric structure of a Riemannian manifold. Typically, the Laplace-Beltrami operator-based manifold regularization can be approximated empirically by the Laplacian regularization associated with the entire training data and its corresponding graph Laplacian matrix. However, the graph Laplacian matrix depends heavily on the prespecified similarity metric and may lead to inappropriate penalties when dealing with redundant or noisy input variables. To address the above issues, this paper proposes a new Semi-Supervised Meta Additive Model (S$^2$MAM) based on a bilevel optimization scheme that automatically identifies informative variables, updates the similarity matrix, and simultaneously achieves interpretable predictions. Theoretical guarantees are provided for S$^2$MAM, including the computing convergence and the statistical generalization bound. Experimental assessments across 4 synthetic and 12 real-world datasets, with varying levels and categories of corruption, validate the robustness and interpretability of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。