提出区间条件风险价值模型,提升非线性回归在扰动与数据污染下的鲁棒性。
Statistical Robustness of Interval CVaR Based Regression Models under Perturbation and Contamination
- 基于区间条件风险价值(In-CVaR)构建非线性回归框架,通过裁剪极端损失增强鲁棒性。
- 理论证明:在多种损失函数下,该模型对数据污染的分布断裂点具有统一上界。
- 首次揭示其在扰动下定性鲁棒性的充要条件,适合高可靠性建模场景。
扰动与数据污染下的鲁棒性是统计学习的核心挑战。本文研究基于区间条件风险价值(In-CVaR)的非线性回归模型,该方法通过裁剪极端损失来增强鲁棒性。尽管已有研究表明In-CVaR在稳健性上优于传统模型,但其在非线性回归中的理论分析仍不充分。本文系统研究了包含线性、分段仿射及带ℓ₁、ℓ₂和Huber损失的神经网络模型在数据污染下的分布断裂点性质,给出了统一的鲁棒性量化结果。同时分析了模型在扰动下的定性鲁棒性:在若干弱假设下,其在Prokhorov度量意义下是定性鲁棒的,当且仅当最大比例的损失被裁剪。本研究从理论与实验两方面验证了In-CVaR在稳健回归中相较于条件风险价值和期望的优势。
原文摘要 · Abstract (English)
Robustness under perturbation and contamination is a prominent issue in statistical learning. We address the robust nonlinear regression based on the so-called interval conditional value-at-risk (In-CVaR), which is introduced to enhance robustness by trimming extreme losses. While recent literature shows that the In-CVaR based statistical learning exhibits superior robustness performance than classical robust regression models, its theoretical robustness analysis for nonlinear regression remains largely unexplored. We rigorously quantify robustness under contamination, with a unified study of distributional breakdown point for a broad class of regression models, including linear, piecewise affine and neural network models with $\ell_1$, $\ell_2$ and Huber losses. Moreover, we analyze the qualitative robustness of the In-CVaR based estimator under perturbation. We show that under several minor assumptions, the In-CVaR based estimator is qualitatively robust in terms of the Prokhorov metric if and only if the largest portion of losses is trimmed. Overall, this study analyzes robustness properties of In-CVaR based nonlinear regression models under both perturbation and contamination, which illustrates the advantages of In-CVaR risk measure over conditional value-at-risk and expectation for robust regression in both theory and numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。