arXiv:2507.10718cs.LGcs.DS2025-07被引 3

同时应对数据污染和分布偏移,提升模型鲁棒性。

Distributionally Robust Optimization with Adversarial Data Contamination

  • 构建新框架,统一处理数据污染与分布变化问题
  • 在数据污染率ε下,估计误差为O(√ε)
  • 适合高可靠性要求的金融、医疗场景

分布鲁棒优化(DRO)为分布不确定性下的决策提供框架,但训练数据中的异常值会削弱其效果。本文提出一种统一方法,同时应对训练数据中的对抗性污染(占比ε)与分布偏移问题。针对广义线性模型及凸Lipschitz损失函数,优化Wasserstein-1 DRO目标。提出新型建模框架,并设计受鲁棒统计启发的高效算法。在有界协方差假设下,证明所提方法仅用被污染数据即可实现对真实DRO目标值的$O( oot{2}{ ext{ε}})$估计误差。本工作首次在高效计算支持下,为双重挑战(数据污染与分布偏移)提供严格理论保证。

原文摘要 · Abstract (English)

Distributionally Robust Optimization (DRO) provides a framework for decision-making under distributional uncertainty, yet its effectiveness can be compromised by outliers in the training data. This paper introduces a principled approach to simultaneously address both challenges. We focus on optimizing Wasserstein-1 DRO objectives for generalized linear models with convex Lipschitz loss functions, where an $ε$-fraction of the training data is adversarially corrupted. Our primary contribution lies in a novel modeling framework that integrates robustness against training data contamination with robustness against distributional shifts, alongside an efficient algorithm inspired by robust statistics to solve the resulting optimization problem. We prove that our method achieves an estimation error of $O(\sqrtε)$ for the true DRO objective value using only the contaminated data under the bounded covariance assumption. This work establishes the first rigorous guarantees, supported by efficient computation, for learning under the dual challenges of data contamination and distributional shifts.

分布鲁棒优化数据污染鲁棒学习广义线性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。