arXiv:2602.18053stat.MLcs.LG2026-02被引 2

揭示了极端尾部数据下CVaR学习的泛化与鲁棒性边界。

On the Generalization and Robustness in Conditional Value-at-Risk

  • 基于极值理论建立CVaR的统一误差展开,分离出阈值敏感项。
  • 在重尾和污染数据下实现最优高概率泛化界,涵盖无限假设类。
  • 提出截断中位数均值估计器,对抗恶意干扰仍保持最优率。

条件风险价值(CVaR)是一种广泛用于罕见但高影响损失场景下的风险敏感目标,但在重尾数据下的统计行为仍不明确。与基于期望的风险不同,CVaR依赖于内生、数据相关的分位数,将尾部平均与阈值估计耦合,从根本上改变了泛化与鲁棒性性质。本文对重尾及污染数据下基于CVaR的经验风险最小化进行了学习理论分析,建立了在最弱矩假设下的紧致高概率泛化与超出风险界,覆盖固定假设、有限与无限假设类,并扩展至β-混合依赖数据;进一步证明这些速率是极小极大最优的。为捕捉CVaR的内在分位数敏感性,推导出统一的Bahadur-Kiefer型展开,分离出均值风险优化中不存在的、由阈值驱动的误差项,这在重尾情形中至关重要。通过提出截断中位数均值法的CVaR估计器,给出了鲁棒性保证,在对抗性污染下仍能达到最优速率。最后,我们证明即使总体最优解分离良好,CVaR决策本身在重尾下也可能固有不稳定。整体结果为何时CVaR学习能泛化且鲁棒,以及何时因尾部稀缺导致不稳定性不可避免,提供了原则性刻画。

原文摘要 · Abstract (English)

Conditional Value-at-Risk (CVaR) is a widely used risk-sensitive objective for learning under rare but high-impact losses, yet its statistical behavior under heavy-tailed data remains poorly understood. Unlike expectation-based risk, CVaR depends on an endogenous, data-dependent quantile, which couples tail averaging with threshold estimation and fundamentally alters both generalization and robustness properties. In this work, we develop a learning-theoretic analysis of CVaR-based empirical risk minimization under heavy-tailed and contaminated data. We establish sharp, high-probability generalization and excess risk bounds under minimal moment assumptions, covering fixed hypotheses, finite and infinite classes, and extending to $β$-mixing dependent data; we further show that these rates are minimax optimal. To capture the intrinsic quantile sensitivity of CVaR, we derive a uniform Bahadur-Kiefer type expansion that isolates a threshold-driven error term absent in mean-risk ERM and essential in heavy-tailed regimes. We complement these results with robustness guarantees by proposing a truncated median-of-means CVaR estimator that achieves optimal rates under adversarial contamination. Finally, we show that CVaR decisions themselves can be intrinsically unstable under heavy tails, establishing a fundamental limitation on decision robustness even when the population optimum is well separated. Together, our results provide a principled characterization of when CVaR learning generalizes and is robust, and when instability is unavoidable due to tail scarcity.

风险建模重尾数据鲁棒学习泛化界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。