arXiv:2502.02710stat.MLcs.LG2025-02NeurIPS被引 5

提出部分可识别下的最优鲁棒性边界,突破传统方法局限。

Achievable distributional robustness when the robust risk is only partially identified

  • 定义新鲁棒性度量:最坏情况稳健风险,适用于不可完全识别场景
  • 实证显示现有方法在未知环境数据增多时性能下降,新方法更稳定
  • 适合安全关键领域中分布偏移不确定的模型设计者

在安全关键应用中,机器学习模型需在最差分布偏移下仍保持良好泛化能力,即具有较小的稳健风险。基于不变性的算法在训练分布异质性强、可识别稳健风险时可证明取得优势。然而实践中此类可识别条件很少满足——这一情形在理论研究中长期被忽视。本文旨在填补空白,研究稳健风险仅部分可识别的更一般设置。我们引入‘最坏情况稳健风险’作为新鲁棒性度量,其始终有定义,且最小值对应算法无关的(总体)极小极大量,衡量部分可识别下的最佳可达鲁棒性。虽可推广至更广场景,本文以线性模型为例明确推导。首先证明现有稳健性方法在部分可识别情况下必然次优。随后在真实基因表达数据上评估,发现随着未见环境数据占比上升,现有方法测试误差持续恶化,而考虑部分可识别性则实现更好泛化。

原文摘要 · Abstract (English)

In safety-critical applications, machine learning models should generalize well under worst-case distribution shifts, that is, have a small robust risk. Invariance-based algorithms can provably take advantage of structural assumptions on the shifts when the training distributions are heterogeneous enough to identify the robust risk. However, in practice, such identifiability conditions are rarely satisfied -- a scenario so far underexplored in the theoretical literature. In this paper, we aim to fill the gap and propose to study the more general setting when the robust risk is only partially identifiable. In particular, we introduce the worst-case robust risk as a new measure of robustness that is always well-defined regardless of identifiability. Its minimum corresponds to an algorithm-independent (population) minimax quantity that measures the best achievable robustness under partial identifiability. While these concepts can be defined more broadly, in this paper we introduce and derive them explicitly for a linear model for concreteness of the presentation. First, we show that existing robustness methods are provably suboptimal in the partially identifiable case. We then evaluate these methods and the minimizer of the (empirical) worst-case robust risk on real-world gene expression data and find a similar trend: the test error of existing robustness methods grows increasingly suboptimal as the fraction of data from unseen environments increases, whereas accounting for partial identifiability allows for better generalization.

鲁棒学习分布外泛化不变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。