用因果框架发现数据缺陷,让机器学习系统更抗压。
Stress-Testing ML Pipelines with Adversarial Data Corruption
- 构建依赖图和污染模板,系统化生成真实数据缺陷
- 仅5%的结构化污染就大幅降低模型性能
- 适合监管评估、模型鲁棒性测试和数据流程优化
结构化数据质量问题,如与人口特征相关的缺失值、文化偏见标签或系统性选择偏差,持续损害机器学习流水线的可靠性。监管机构日益要求高风险系统能抵御这类现实且相互关联的错误,但当前鲁棒性评估多采用随机或过于简单的污染方式,未覆盖最坏情况。本文提出SAVAGE,一个基于因果启发的框架:(i)通过依赖图和灵活污染模板形式化建模真实数据问题;(ii)系统性发现使目标性能指标下降最大的污染模式。SAVAGE采用双层优化,高效识别脆弱数据子群体并调优污染强度,将整个机器学习流水线(含预处理及可能不可微模型)视为黑箱。在多个数据集和任务(数据清洗、公平学习、不确定性量化)上的实验表明,即使仅约5%的结构化污染,也显著损害模型性能,远超随机或人工设计的错误,颠覆现有技术的核心假设。因此,SAVAGE提供了一种实用的流水线压力测试工具,可作为鲁棒性方法评估基准,并为设计更稳健的数据工作流提供行动指导。
原文摘要 · Abstract (English)
Structured data-quality issues, such as missing values correlated with demographics, culturally biased labels, or systemic selection biases, routinely degrade the reliability of machine-learning pipelines. Regulators now increasingly demand evidence that high-stakes systems can withstand these realistic, interdependent errors, yet current robustness evaluations typically use random or overly simplistic corruptions, leaving worst-case scenarios unexplored. We introduce SAVAGE, a causally inspired framework that (i) formally models realistic data-quality issues through dependency graphs and flexible corruption templates, and (ii) systematically discovers corruption patterns that maximally degrade a target performance metric. SAVAGE employs a bi-level optimization approach to efficiently identify vulnerable data subpopulations and fine-tune corruption severity, treating the full ML pipeline, including preprocessing and potentially non-differentiable models, as a black box. Extensive experiments across multiple datasets and ML tasks (data cleaning, fairness-aware learning, uncertainty quantification) demonstrate that even a small fraction (around 5 %) of structured corruptions identified by SAVAGE severely impacts model performance, far exceeding random or manually crafted errors, and invalidating core assumptions of existing techniques. Thus, SAVAGE provides a practical tool for rigorous pipeline stress-testing, a benchmark for evaluating robustness methods, and actionable guidance for designing more resilient data workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。