arXiv:2603.15867cs.LG2026-03

用数学约束生成真实扰动,检测模型对数据变化的脆弱性。

Evaluating Black-Box Vulnerabilities with Wasserstein-Constrained Data Perturbations

  • 基于最优传输约束特征统计量,生成语义一致的扰动
  • 在真实数据集上验证,提供可解释的鲁棒性诊断结果
  • 适用于表格和图像数据,无需依赖具体模型

机器学习工具的广泛应用带来了可解释性不足等关键挑战。本文提出一种全局可解释性框架,利用最优传输与分布鲁棒优化分析机器学习算法对受限数据扰动的响应。该方法对特征级统计量(如亮度、年龄分布)施加约束,生成保留语义结构的真实扰动。我们构建了一个模型无关的诊断基准,适用于表格与图像领域,并具备坚实的理论保障。在真实数据集上的验证表明,该方法能提供补充标准评估与公平性审计工具的可解释鲁棒性诊断。

原文摘要 · Abstract (English)

The growing use of Machine Learning (ML) tools comes with critical challenges, such as limited model explainability. We propose a global explainability framework that leverages Optimal Transport and Distributionally Robust Optimization to analyze how ML algorithms respond to constrained data perturbations. Our approach enforces constraints on feature-level statistics (e.g., brightness, age distribution), generating realistic perturbations that preserve semantic structure. We provide a model-agnostic diagnostic bench that applies to both tabular and image domains with solid theoretical guarantees. We validate the approach on real-world datasets providing interpretable robustness diagnostics that complement standard evaluation and fairness auditing tools.

可解释性鲁棒性最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。